This is not a model writing a summary. Every topic is produced by an agent-orchestrated pipeline: plan the retrieval first, gather evidence chain by chain, attach a mandatory retrieval log to every research document, fail the mechanical validator and it goes straight back, and clear a global deduplication gate before anything is admitted.
The gates are mechanical —
a script decides, a failure means exit 1, and the output does not enter
the registry. There is no “close enough, let it through”.
An agent will naturally smooth things over. So we do not rely on good intentions — we rely on countable floors: a finished document must hand over its retrieval ledger, and one that falls short is sent back to be rerun.
Written into the agent prompt, mechanically enforced by script.
| # | Query | Hits | Key hit | Used in | Note |
|---|---|---|---|---|---|
| 1 | Anderson 2018 "On Evaluation of Embodied Navigation Agents" SPL formula definition | 5 | paper notes | §1, §10 | mostly Chinese results |
| 3 | arxiv 1807.06757 SPL formula geodesic distance | 5 | paper notes | §1, §10 | paper identity confirmed |
| 7 | aihabitat.org challenge PointNav ObjectNav evaluation rules success SPL threshold | 1 | no direct hit | — | search failed |
| 8 | "SPL" "Success weighted by Path Length" formula "S * l" geodesic shortest path | 5 | contains SPL formula | §1, §8 | formula confirmed |
| 11 | "soft-SPL" habitat definition formula "distance_to_goal" "start_distance" | 5 | soft_spl found in nav.py | §5, §8 | SoftSPL confirmed |
85 seconds, unedited and not sped up.
The files that appear in the recording, published as they are and unedited. These are frozen snapshots — however the pipeline changes later, the versions here do not move. The research documents are in Chinese, the source language of the compiled corpus.
The document on the left of the recording — the entry point for batch compilation, laying down the unattended-operation rules and the Phase 0–4 execution flow.
The all-chain research summary for one topic: breadth scan → candidate chain identification → red-team review → per-chain quality gating.
Five per-chain deep-research documents under that topic — section 12, “Retrieval log”, is the full version of the ledger above: