Changelog

What shipped on reranker.uk — demo improvements, new guides, and site infrastructure.

28 Sep 2026 — transformers.js 4.3.0, and what report-only caught

  • The demo runs on transformers.js 4.3.0 — up from 3.5.1, with ONNX Runtime Web 1.22 → 1.31. Verified against live jsDelivr and HuggingFace on all three demo models before merging, not just checked against the source
  • The report-only CSP earned its keep — its first run against real traffic flagged that HuggingFace redirects model weights to separate CDN hosts the policy didn't list. Enforcing, that would have blocked every model download in the demo. Fixed before it was ever enforced
  • …and caught the upgrade too — transformers.js 4 loads part of ONNX Runtime from a blob: URL by default. The demo now switches that cache off instead of loosening the policy for every page
  • A stricter real-network test — the daily check now runs every model in the demo's picker, not just the default, and fails if an obviously relevant passage doesn't outrank an unrelated one, so a model that loads but ranks wrong no longer passes

27 Sep 2026 — Security headers, report-only for now

  • public/_headers added — X-Content-Type-Options, X-Frame-Options, Referrer-Policy, Permissions-Policy, and a Content-Security-Policy, generated at build time rather than hand-maintained
  • CSP ships report-only — this repo's sandboxed dev environment can't reach real jsDelivr/HuggingFace/hf-mirror.com traffic or real WASM instantiation to confirm a stricter policy wouldn't break the demo in practice, so it logs violations instead of blocking anything until that's been checked against the real thing
  • No hand-copied hash — the CSP's script-src allows the site's one inline script (early theme detection) via a sha256- hash computed from its actual content at build time, not typed in by hand where it could drift
  • Both smoke tests now watch for CSP violations — tests/helpers/csp.mjs, asserted in both the mocked and real-network demo tests; verified the mechanism actually catches something by deliberately breaking style-src locally and watching the test fail

26 Sep 2026 — ms-marco MiniLM's BEIR figure was never published

  • ms-marco MiniLM-L6 now reads "not published" — this table carried it at an unsourced ~55.0; sentence-transformers' own docs report NDCG@10 on TREC DL 19 (74.30) and MRR@10 on MS MARCO Dev (39.01) for this model, not a BEIR average. Found while checking the rest of the table after the mxbai correction below — same failure shape, different model

26 Sep 2026 — mxbai-rerank-v2, a corrected BEIR figure, and a demo smoke test

  • mxbai-rerank-large-v1's BEIR score corrected — this table carried an unsourced ~62.1 for months; mixedbread's own comparison table puts it at 49.32*. A reader flagged the gap between that number and mixedbread's own figures for their newer models, which is what surfaced it
  • mxbai-rerank-v2 added — a newer, Qwen2.5-based, RL-trained generation (base-v2 0.5B, large-v2 1.5B) that adds 100+ languages including Chinese; v1 was English-only, which the table didn't make clear
  • data/models.json — benchmark numbers that have been through a correction like the one above now live in one sourced file instead of a hand-typed literal in every page that mentions them; the build refuses to render one with no source and warns when one hasn't been re-checked in 6 months
  • A smoke test for the live demo — npm test runs the demo's load → score → render pipeline against a fake model on every PR; a separate daily job runs it against the real jsDelivr / HuggingFace chain, which is what would actually catch one of those going dark

18 Sep 2026 — URLs without the .html, and two pricing corrections

  • URLs dropped their .html — every internal link, the sitemap, hreflang tags and JSON-LD now point straight at the extensionless URL Cloudflare was already redirecting to, cutting a 307 round-trip off every internal navigation
  • Voyage's free-tier grant moved on — the 200M free-token allowance is no longer on rerank-2.5; it now ships with the rerank-3 preview, added to the model table and cost calculator alongside rerank-3-lite
  • Cost calculator token formula fixed — Voyage bills the query once per document reranked, not once per query; the calculator undercounted tokens (and cost) on any workload with more than one candidate per query
  • Cohere chunking, clarified — an outside audit suggested Cohere silently splits passages over 500 tokens into extra billable chunks; that isn't how it works (context is 32,768 tokens and chunking is opt-in), so the calculator page now says so directly instead of adopting the wrong model
  • Homepage diagram, now bilingual — the "Reranking in one diagram" figure was inline SVG the translation pass couldn't see; its labels and caption now render in Chinese on /zh/
  • Model table accessibility — the comparison table gets a caption, column/row headers, and a scrollable, keyboard-focusable region for screen readers and narrow viewports

25 Aug 2026 — Two guides for the questions people actually ask

  • Reranking didn't help — a diagnostic walkthrough of the seven reasons a rerank stage shows no lift, starting with the one that explains most of them: retrieval never returned the right document
  • Rerank on a vector database — the retrieve-wide-then-rerank pattern with code for pgvector, Qdrant and Elasticsearch, including the HNSW ef_search trap that makes a wider limit return padding instead of candidates
  • /llms.txt — a generated map of the site for assistants that read one before citing a source, including how to read our benchmark footnotes
  • Feed autodiscovery — the changelog RSS is now advertised from every page, not just the changelog

14 Aug 2026 — Cost calculator, and demo links that land somewhere

  • Rerank cost calculator — Cohere bills per search, Voyage per token, so the cheaper vendor flips with passage length and top-k. Put your own volume in and see where the line sits
  • Scenario deep links — all 25 demo links across the guides and model pages now open the scenario the page is actually about, via a new ?s= parameter
  • Shorter share links — an untouched built-in scenario shares as ?s=rag instead of a 900-character URL
  • Translations survive link edits — the build now retargets links inside a translation instead of dropping it back to English

11 Aug 2026 — Real Chinese URLs, and a 2026 model refresh

  • Chinese lives at /zh/ — every page is now pre-rendered in Chinese at its own URL with a self-referencing canonical, instead of a client-side toggle that left all three hreflang tags pointing at one page
  • Lighter pages — translation moved to build time, so ~200 KB of dictionaries no longer ship to the browser
  • Sitemap is generated — built from the page tree with hreflang alternates and lastmod from git, replacing a hand-maintained file that had drifted
  • Cohere Rerank 4 — rerank-v4.0-pro and rerank-v4.0-fast replace v3.5; 32k context, billed per search
  • Voyage rerank-2.5 — 32k context and instruction following; the index table's per-doc pricing was wrong and is now per token
  • Jina v3 numbers firmed up — 0.6B, 61.94 BEIR nDCG@10, 64 docs in a 131K context
  • Honest Qwen3 rows — dropped an unverifiable "~75+" figure for the 8B; reports put the 4B slightly ahead of it
  • Design refresh — fluid type scale, a real elevation ramp, and one focus-visible treatment site-wide

25 Jun 2026 — Model landscape refresh

  • Qwen3-Reranker — 0.6B / 4B / 8B rows + deep review page
  • Jina v3 — listwise flagship; tiny kept for browser demo
  • Table adds — gte-reranker-modernbert-base, NVIDIA nv-rerankqa / Nemotron
  • Self-host / homepage / chooser — GPU default Qwen3-4B; bge remains CPU path
  • Score footnotes — MTEB-R* vs classic BEIR; next review Oct 2026

25 Jun 2026 — Minimal polish pack

  • Honest benchmarks — emerging rows use ≈ / n/a; next review Sep 2026; pricing lag disclaimer
  • Demo max_length — 256 / 384 / 512 tokens; char warnings follow selection
  • Lazy transformers.js — loaded on first Rerank only
  • Instruction-rerank guide — task-shaped ranking
  • Homepage — beyond five families + links to ColBERT / instruction guides

25 Jun 2026 — Content expansion & polish

  • Models table — architecture column; Qwen3-Reranker, Contextual AI, ColBERTv2, ms-marco browser rows
  • Late-interaction guide — ColBERT & when to skip cross-encoder rerank
  • Demo presets — E-commerce + Multilingual (7 scenarios); ms-marco ?m= on models table
  • JSON-LD — inLanguage follows zh/en toggle
  • Changelog RSS — /changelog.rss feed

25 Jun 2026 — i18n & SEO completion

  • Model page i18n — pills, TOC, meta dates, Pros/Cons, Other models for all five families
  • Changelog + Privacy i18n — full Chinese body on both pages
  • hreflang — en, zh-Hans, x-default on every page (same URL, client-side toggle)
  • og:locale:alternate — swaps with primary locale on language toggle

25 Jun 2026 — Low-priority polish

  • Guide i18n — full Chinese body for self-host, scenario, hybrid retrieval, and evaluate guides
  • Compressed share links — demo ?z= gzip when URLs exceed ~1600 chars
  • Preset mobile layout — 2-column grid on narrow screens
  • og:locale — switches to zh_CN when language toggle is 中文
  • Dual-diff a11y — table caption, row headers, empty state, aria-labelledby

23 Jun 2026 — Medium-priority UX & content

  • Self-host guide — sentence-transformers, serving, ops
  • Scenario guide — RAG vs support vs legal vs code
  • Passage char counts — per-list stats + 512-char truncation warning
  • New presets — Technical docs, Code search
  • CSV export — Copy CSV alongside JSON and Markdown
  • Light theme — toggle in nav, persisted in localStorage

23 Jun 2026 — Demo UX round 2

  • Model loading panel — progress %, ETA, file name, cache status
  • JSON passages — paste a JSON array or { "passages": [...] } object
  • Dual-model diff view — aligned table with score and rank deltas
  • Models table — sort, filter, and jump to demo with ?m=
  • Mobile passage editor — add/remove list on small screens
  • Error hints — classified messages for network, WebGPU, memory, and limits
  • Changelog + Privacy — this page and a short privacy statement
  • Nav — Home and Guides links in the top bar

23 Jun 2026 — Full release (plan items 1–5, 7–10)

  • Build system: src/partials + src/pages → scripts/build.mjs
  • Demo: three-column results, bi-encoder proxy, dual-model compare, WebGPU, URL sharing
  • Guides index, hybrid retrieval, evaluate rerankers
  • Models: Last verified June 2026, chooser cards, mxbai on homepage
  • i18n: data-i18n keys + shared dictionary; EN/中文 toggle
  • Footer GitHub link; sticky nav; aria-live status

21 Jun 2026 — Initial launch

  • Educational guides on rerankers, cross- vs bi-encoder, and RAG
  • Model comparison pages for bge, Cohere, Jina, Voyage, mxbai
  • In-browser cross-encoder demo with transformers.js
  • Deployed on Cloudflare Workers (static assets)

Try the demo →