Getting started
Benchmark first
Start with one slice each from BEIR, FEVER, and LongMemEval. They expose retrieval, evidence, updates, temporal reasoning, and abstention. Record quality, latency, tokens, storage, and corpus revision from the first prototype; the same cases become regression gates as components change.