Enterprise AI Bootcamp Demo 1

The model was not the problem

Every retrieval, rerank and generation emits a span using the OpenTelemetry GenAI attribute names, so the question "was this the model or the retrieval?" is answered by reading the trace rather than by arguing about it.

Trace 22b2a0c0d80f4d30

eval.structure · 7 ms · status ok

Question
What is the air intake filter service interval for the Davey 6M125 125 CFM diesel compressor?
Mode
structure
Documents returned
Answer refused
Cost
$0.00000

Span tree

retrieval retrieve
7.1 ms
9 attributes
atlas.doc_ids ['tm-9-4310-241-15', 'tm-5-4310-221-15', 'tm-5-4310-277-14', 'tm-5-4310-275-14', 'tm-5-4310-354-14']
atlas.families ['A', 'B', 'C', 'D']
atlas.k 5
atlas.mode structure
atlas.n_results 5
atlas.pages [70, 26, 18, 34, 195]
atlas.reranked False
atlas.top_score 0.0164
atlas.withheld 0

The counterfactual: same model, correct passages

The argument that the bottleneck is retrieval is only worth making if it can be falsified. This runs the identical question through the identical model and the identical prompt, changing one thing: the second run is handed the passages the gold set says contain the answer. If the answer becomes correct, the model was never the constraint. The gold passages are the ones pipeline/questions.py located in the corpus, not passages chosen to make the point.

Index built by scripts/build_index.py at 2026-08-10 01:44:34 in 711s: 25 documents, 3,146 pages, 12,590 improved chunks / 5,434 naive chunks. Embeddings: Snowflake/snowflake-arctic-embed-s (int8 ONNX, Apache-2.0). Reranker: cross-encoder/ms-marco-MiniLM-L-6-v2 (ONNX, Apache-2.0).