Waterloo pre-trains a 3B reranker from scratch, no open-weight backbone, beats GPT-6.1 on TREC DL19

Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search

Jimmy Lin, Sahel Sharifymoghaddam, Lingwei Gu, Nour Jedidi

cs.IR, cs.CL

2026-10-08

Waterloo pre-trains a 3B model from scratch, no third-party backbone, and fine-tunes it into a reranker that beats every fine-tuned baseline: 0.641 nDCG@10 on TREC DL, 0.547 on BEIR.

What problem this solves

The default way to build a reranker is to fine-tune an open-weight backbone like Qwen or Gemma on relevance annotations. Convenient, but the supply chain belongs to someone else: the pre-training corpus, data mix, and alignment pipeline behind that backbone are all secret. You get weights, not provenance. There is no way to audit data licensing, screen for baked-in bias or backdoors, or rebuild the model on a different data mix. The paper cites the June 2026 episode in which Anthropic suspended Fable 5 access under a US government directive (later lifted) as a reminder that losing a critical capability to someone else's decision is not hypothetical.

Project Greenhouse, from Jimmy Lin's group at Waterloo, tests one thesis: with public data and modest compute, can you train competitive search models end to end, owning every dependency? The first milestone is a pointwise reranker, chosen for practical reasons: a query-document pair goes in, a scalar score comes out, so iteration is fast; it works as a standalone component in any retrieval pipeline; and it is a stepping stone toward agentic search, usable as a relevance judge or an RL reward signal.

Method

Two steps, with no instruction tuning, alignment, or mid/post-training in between, and no teachers or synthetic data.

Pre-training from scratch. The backbone copies Karpathy's nanochat depth-34 configuration: a 34-layer decoder-only transformer with a 2,048-token context, 3.29B parameters total. 1.21B of those sit in value-embedding lookup tables that cost almost no compute, so per-token FLOPs are close to a dense 2B model. Data is the public ClimbMix corpus, one full pass, 287B tokens. The optimizer setup follows nanochat (Muon for weight matrices, AdamW for embeddings); the only real change is the batch schedule, which starts at 262K tokens and doubles three times to about 2.1M while the peak learning rate stays fixed. Early in training the model is weak and small batches already point in a useful direction, which saves data; later, as batch signal-to-noise drops, growing the batch itself acts as learning-rate decay. Hardware: one 8xH100 server, FP8 matmuls, roughly 1,600 GPU-hours end to end.

Supervised fine-tuning. The prompt is 'Query: {query} Passage: {passage} Relevant:', the score is the logit difference between the next-token candidates ' true' and ' false', and all other output rows are dropped, leaving 3.22B parameters. Training data is RLHN-250K, about 247K query-document relevance pairs. The loss is LCE, a softmax over one positive and K hard negatives for the same query. One epoch, single-GPU scale. The same recipe ran four times, and the four checkpoints were weight-averaged into a model soup for release.

The paper also formalizes LLM training as a directed property hypergraph: vertices are artifacts (corpora, weights), hyperedges are computations (recipes plus configs). A model is 'fully open and sovereign' if and only if every vertex and edge upstream of it is public. By that standard, anything fine-tuned from a backbone with an undisclosed pre-training process fails.

Results

nDCG@10, all models reranking top-100 BM25 candidates:

MethodType / sizeTREC DL meanBEIR mean
BM25first-stage baseline0.3930.398
Gaggle (this paper)pointwise 3B0.6410.547
MonoT5pointwise 3B0.6040.520
Qwen3-Rerankerpointwise 4B/8B0.615 / 0.6050.541 / 0.541
RLHN Qwen2.5pointwise 3B0.6150.520
Jina-Reranker-v3listwise 0.6B0.6110.511
Gemma-4listwise 26B, zero-shot0.6370.548
Qwen3.8listwise 27B, zero-shot0.6330.554
GPT-6.1 Sollistwise frontier0.6530.572
Oraclereranking upper bound0.7890.735

Gaggle beats BM25 on all 12 collections, and its means top every fine-tuned baseline on both benchmarks. On TREC DL only GPT-6.1 Sol is higher; on BEIR it trails Sol and Qwen3.8 and ties Gemma-4. On DL19 alone it scores 0.757, the best non-oracle number in the table, above Sol's 0.750, and it also beats Sol on TREC-COVID. The four training runs establish a noise scale: about 0.001 standard deviation on benchmark means, 0.003 per collection; smaller gaps count as noise. The soup's margin over the average of its four ingredients (+0.001/+0.003) sits right at that edge.

Two side findings carry more signal than the headline. First, the backbone is a weak general model: CORE 0.458 and MMLU 0.493, below nearly every model in the paper's comparison table, most of them trained on 1.7T to 34T tokens. It is still enough for reranking. Second, stronger general backbones fine-tune into worse rerankers under the shared recipe (Gemma-4-E2B: BEIR 0.537 vs. 0.544). The diagnosis is weight scale: AdamW-pretrained backbones have weight matrices roughly five times smaller than this Muon-pretrained one, so the 5e-5 peak that works here pushes their loss up; at 1e-5, MiniCPM5-2B reaches BEIR 0.551 and overtakes Gaggle. Backbone comparisons under a shared recipe mislead unless the learning rate follows the weight scale. On the data side, training only on MS MARCO drops BEIR to 0.470; the multi-source mix is what carries out-of-domain transfer.

Why it matters

This is an existence proof: no third-party backbone, no distillation, one 8xH100 server for about 200 hours plus single-GPU fine-tuning buys a reranker in the same band as zero-shot 26B/27B listwise rerankers and the purpose-built Qwen3-Reranker. That is replicable by an academic lab. Every artifact along the way (corpus, code, configs, checkpoints) is public; if a slice of training data turns out to be a problem, you delete it and retrain, an option open-weight models do not give you. Training from scratch also makes controlled interventions on pre-training data possible, which is the only way to study where a model's knowledge actually comes from.

The cold water is equally clear. The authors frame this as an existence proof, not an attribution study. MiniCPM with a retuned learning rate beats it on BEIR, so from-scratch is the sovereign path, not the performance-optimal one. The 2,048-token context also limits direct use on long documents.

Limitations

Acknowledged by the authors: comparison models differ in size, data, and training and inference procedures, so no single design choice can be credited; no hyperparameter search was done for the baseline or the contrastive conditions; the scope is pointwise reranking only, far from agentic search; and the model does not compete on general capability.

Concerns from reading: three of the seven BEIR collections (SciFact, SCIDOCS, FiQA) overlap with the RLHN-250K training sources, and the paper counts only the other four as out-of-domain, so the BEIR mean includes in-domain scores (the RLHN baseline shares the same data family, so the comparison remains fair). 'No teacher, no synthetic data' deserves a footnote: RLHN-250K's false negatives were relabeled by LLMs during dataset construction, and the paper openly declines to settle where that falls on the openness spectrum. The soup's gain is at noise level. And 1,600 GPU-hours is modest by industry standards but still a budget line for most labs; 'a handful of GPUs' is relative.

Terms

Source

What people are saying

Related papers

All paper explainers