Revision history for Cavil::Matcher 1.05 2026-09-03 - Removed the on-disk fingerprint index (fp_build/fp_open, the Cavil::Matcher::Fingerprints package, and Cavil::Matcher::FpIndex); consumers now keep the searchable index in their own database. The fingerprint_file and content_hash primitives are unchanged. 1.04 2026-08-27 - FpIndex now caches opened segments and reuses them across searches instead of reopening every segment on every query. Opening a segment is a fixed per-query cost that otherwise dominated search time on a long-lived query server, regardless of the query. Segments are immutable append-only files, so a cached handle always maps the same bytes; a handle whose segment has been compacted out of the manifest is dropped on the next search. Results are unchanged. - score() and FpIndex::search() take an optional minimum containment; matches below it are dropped inside the scorer instead of being returned for the caller to filter. A query dominated by common fingerprints can match hundreds of thousands of contents, almost all coincidental, so applying the floor at this level avoids building and returning them. Defaults to 0 (return everything). - score() and FpIndex::search() also take an optional max_df: query fingerprints appearing in more than that many records of a segment are ignored (boilerplate that matches nearly everything, adding only coincidental hits and skewing containment). Off by default (0). Intended for a compacted single-segment index, where a segment's record count for a fingerprint is its document frequency. 1.03 2026-08-27 - Fingerprint scoring now reports, for each matched fingerprint, the query fingerprint value it was, appended to each region tuple after the existing start line and span. This lets a caller map matches back to their query positions - enough to tell an aligned copy from scattered coincidental hits, and to derive each query fingerprint's document frequency from a full result set. Backward compatible: the (start_line, span) pair at the front of every region tuple is unchanged. 1.02 2026-08-27 - Snippet code search: a new winnowed-fingerprint index for "where does this code already exist" provenance queries, alongside the license-pattern matcher. fingerprint_file winnows a file's tokens (k tokens per gram, w grams per window) into content fingerprints; a fingerprint segment stores them keyed by a 128-bit content hash rather than filenames (so it stays lean and joins to the Cavil database), and scoring ranks candidates by both-direction containment with exact matched-line spans for highlighting. - Cavil::Matcher::FpIndex is the pure-Perl lifecycle over a directory of fingerprint segments (incremental add_segment, manifest/generation, memory-mapped shared reads), reusing the same on-disk discipline as the pattern index: versioned, CRC-checksummed, structure-validated on every open, with corrupt or foreign files rejected rather than mis-read. - Byte-identical files are de-duplicated by content hash, so the same source shipped across many package versions is stored once, with no package-name parsing. - Additive only: the license-pattern format is unchanged and no migration is needed. 1.01 2026-08-04 - Big indexing speedup: the scan path (load/attach) no longer re-checksums a segment's whole payload on every open. A compiled segment is an immutable, atomically-published cache that is already CRC-checked when written, so re-CRCing hundreds of megabytes in every forked index job was almost the entire cost of a load (~97% in a 160 MB corpus: ~293ms -> ~8ms). Structural validation still always runs, so bad segments are rejected and scanning stays memory-safe; only per-scan detection of on-disk bit-rot is dropped. - CRC is now verified where it is meaningful: once at publish (dump() fully verifies before the atomic rename) and on demand via the new Engine verify($file) method (an fsck for a cache dir). - Same treatment for the closest-match "bag of patterns" cache (Bag::load, used by analyze): the load path trusts the payload CRC (the record parse still bounds-checks every read), with a new Bag verify($file) method and a publish-time check. - No on-disk format change: existing segments load as-is, no reindex, and old/new engines interoperate on the same files (in-place upgrade is seamless). 1.00 2026-07-22 - Initial release: the next-generation license pattern matcher for Cavil. - Segmented, incrementally-updatable index: adding or removing a pattern no longer rebuilds the whole compiled cache (base + append-only delta segments + tombstones). - Versioned, CRC-checksummed, fully-validated on-disk segment format (no unchecked reinterpret_cast); corrupt or hostile segments are rejected, never mis-read. - Segments are memory-mapped read-only and queried in place, so index workers on a host share one physical copy instead of each materialising its own. - Pure-Perl manifest / segment lifecycle (Cavil::Matcher::Index, ::Manifest); a small, frozen C++ core (tokenizer, scan, SpookyV2 hashing) kept bit-for-bit compatible with Spooky::Patterns::XS, so no database migration is needed to switch engines. - Hardened against arbitrary input (binaries, malformed samples): never crashes. - 100% statement/branch/condition coverage of the Perl surface; developer-only differential tests against the previous engine in xt/.