The model was not the problem
Every retrieval, rerank and generation emits a span using the OpenTelemetry GenAI attribute names, so the question "was this the model or the retrieval?" is answered by reading the trace rather than by arguing about it.
Trace 8d2ff127bce74fb3
eval.improved · 2442 ms · status ok
- Question
- What is the torque spec for the Atlas Copco GA 30 VSD+ gearbox?
- Mode
improved- Documents returned
- Answer refused
- Cost
- $0.00000
Span tree
9 attributes
['tm-5-4310-275-14', 'tm-5-4310-354-14', 'tm-5-4310-354-14', 'tm-3-1040-244-35p', 'tm-5-4310-227-34p']['B', 'C', 'D', 'E']5improved5[97, 20, 129, 6, 37]True-6.88450The counterfactual: same model, correct passages
The argument that the bottleneck is retrieval is only worth making if it can
be falsified. This runs the identical question through the identical model and the identical
prompt, changing one thing: the second run is handed the passages the gold set says contain
the answer. If the answer becomes correct, the model was never the constraint. The gold
passages are the ones pipeline/questions.py located in the corpus, not passages
chosen to make the point.
Index built by scripts/build_index.py at 2026-08-10 01:44:34 in 711s: 25 documents, 3,146 pages, 12,590 improved chunks / 5,434 naive chunks. Embeddings: Snowflake/snowflake-arctic-embed-s (int8 ONNX, Apache-2.0). Reranker: cross-encoder/ms-marco-MiniLM-L-6-v2 (ONNX, Apache-2.0).