Skip to main content
Rail402’s search is scored on a frozen, graded dataset by a harness that runs the production SearchService. It reproduces with one command from a fresh clone. The dataset, method and results are documented in full in tools/search-eval/README.md; this page summarises them.

Reproduce

Prerequisites: Node.js 24.11 or later (.node-version pins 24.19.0), pnpm 11.22 through Corepack (corepack enable), and network access to huggingface.co on the first run, which downloads the embedding model (about 90 MB). A run takes about 15 seconds on a laptop CPU. Every number except latency reproduces exactly on the machine that recorded the baseline; other CPU models have not been checked, and floating-point inference may differ on them.
The pinned embedding model (all-MiniLM-L6-v2, revision 1110a243fdf4706b3f48f1d95db1a4f5529b4d41, hashes in packages/search/models/all-minilm-l6-v2.json) is downloaded from Hugging Face on first use and verified against its SHA-256 hashes. The harness refuses to run on data whose hashes do not match tools/search-eval/data/manifest.json.

Results

Baseline search-eval-v1.1.0, test split, 81 answerable queries. Relevant means grade 2 or higher. nDCG uses linear gain with the ideal ranking over all judged listings (trec_eval semantics).
  • 95% bootstrap intervals for hybrid: nDCG@10 [0.700, 0.780], Recall@20 [0.846, 0.940].
  • p-values come from a paired randomization test.
  • Latency covers the whole search call, with single-threaded model inference, on an Intel Core Ultra 5 225 (environment in tools/search-eval/data/report.json).
  • Filter violations: 0 across 1,216 results of 43 constrained queries, checked by an implementation independent of the service (tools/search-eval/src/constraints.ts).
  • Constraint extraction, against hand-labelled constraints (45 queries): network 15/16, type 23/23, asset 13/14, price 14/15, with no spurious filters.
By query class, all splits, answerable queries:

Filter conformance

Network, asset, scheme, recipient and price filters are each tested on their own, apart from the judged dataset, by tools/search-eval/src/conformance.ts. It generates a catalog of 172 listings built to stress the filters:
  • both networks, and the exact and upto schemes;
  • recipients as a G… account, a muxed M… address on it, another G… account and a contract C…;
  • assets with 6, 7 and 18 decimals, and one the service does not know;
  • prices one base unit below, at and above each ceiling;
  • listings whose options each satisfy part of a filter, which must never match the whole of it.
31 queries (5 network, 8 asset, 3 scheme, 7 recipient, 8 price) use the filters as parameters and as query text, alone and combined. Each runs lexical-only and hybrid and is paged to the end with cursors. An independent check compares the results both ways: a result, or a payment option shown in it, that breaks a filter is a violation; a listing that satisfies every filter but is not returned is a miss. Results are counted over both modes. The gate fails on any violation or miss.

The gate

pnpm eval:check runs in CI and fails when:
  • test nDCG@10 or Recall@20 drops more than 0.02 below tools/search-eval/data/baseline.json;
  • hybrid scores below BM25 on either metric, as measured now or as recorded in the baseline;
  • any result violates an applied filter, or the filter-conformance run finds a violation or a miss;
  • the dataset or the model changed without a re-baseline;
  • tools/search-eval/data/baseline.json differs from the copy at the release tag it names (search-eval-vX.Y.Z), so a change cannot lower the bar by editing it.
A change that improves search re-baselines in a dedicated commit with pnpm eval --write-baseline search-eval-vX.Y.Z, and that commit gets the new annotated tag. CI fetches tags to check the baseline against it.

Dataset

  • Public listings (4) are every resource in Rail402’s hosted testnet catalog when the dataset was built: the public demo seller’s four resources ("source": "public", with the catalog URL and retrieval time). They were read with plain GET requests; the raw capture, with page hashes and the facilitator version, is in tools/search-eval/construction/public/capture.json.
  • Other catalogs are not imported: Rail402 does not harvest other facilitators’ catalogs, whose licensing and provenance it cannot vouch for, so services listed only elsewhere are not included.
  • Sample listings (496) are synthetic and labelled so ("source": "sample"; the 461 web URLs use hosts under the reserved .example domain, and the 35 mcp:// URLs use the host tool). They were drafted by six writers from a taxonomy of 15 categories and 107 capabilities, in four styles, with competing providers, sibling endpoints, multi-tool MCP servers and 10 adversarial listings. Each draft’s discovery metadata is produced with the stock declareDiscoveryExtension and passed through Rail402’s own extractCandidate, so the corpus is exactly what the catalog stores after a settlement. Four samples that no judge graded above 0 were dropped to keep the corpus at 500.
  • Queries were written by three writers who saw only the taxonomy, never the listings, and were frozen in a commit before any system ran on them. A typo query shares its source query’s dev or test split.
  • Judgments pool the results of six systems per query (22 to 58 candidates) and were graded blind by two independent judge sets. The public listings are judged against every query instead of pooled. The sets agreed exactly on 96.8% of pairs and within one grade on all of them (quadratic-weighted kappa 0.965). Grades one apart take the lower. A listing that breaks a query’s labelled constraints is graded 0 by code: 214 grades were overridden this way.

Limitations

  • The sample listings and the queries are synthetic. The corpus is diverse by construction, but it is not a sample of a real catalog.
  • Search has no confidence threshold that yields an empty result: it returned results for 33 of 34 unanswerable queries.
  • The evaluation prices listings in USDC, EURC and XLM on both networks. A deployment’s search knows the assets its facilitator accepts (USDC by default), which are the only assets its catalog can hold listings in.
  • With 12 to 40 queries per class, per-class differences under about 0.07 nDCG are within noise.