This page is the short version, written to be read by someone deciding whether to trust the output. Indexing now covers NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, EMNLP, NAACL, AAAI, IJCAI and arXiv, starting from a deliberately narrow NeurIPS/ICML/ICLR core before widening to the rest on the same pipeline. Five of the six stages that write edges are running at full scale against the corpus today — identity resolution, review ingestion, code-link ingestion and method/concept extraction have all written real edges into the live graph. The narrower one is method-to-method relationship labeling (generalizes / improves on / related to / evaluated on), which is real and corroborated where it exists but has not been run against the full method population yet. Each stage below says which it is.
{{ s.what }}
{{ s.limit }}
Everything above is about the graph's structure. This is about one specific claim on top of it — that a piece of code for a method actually does what the paper says — and it runs as its own process, on its own curated slice, not as a stage every paper reaches.
{{ s.what }}
{{ s.limit }}
Each relation carries where it came from — a structured API, an explicit sentence in the paper, or a model extraction with a confidence score. No edge is written without it. This is the difference between having a graph and having a traceable one.
One node per reviewer per paper, rating and confidence kept raw. Averaging across reviewer confidence is a known, biased statistic, so it is never done. Reviewer identity is a permanent omission — not a field to fill in later.
Each node carries a readiness rollup — dense, sparse, or unresolved — computed from its own provenance mix. A deep trace over a sparse region says so before it runs, instead of returning a confidently wrong answer. The rollup is specified and gated into the retrieval design; it is computed once the edges it reads exist.
The extraction mechanism was proven on a deliberately narrow core — NeurIPS, ICML and ICLR, 2020–2025 — before widening. Vision and language venues followed, and arXiv is indexed on a fixed cadence with its unreviewed status carried on the record rather than dropped.
This list is not comprehensive today and won't claim to be. Publication venues evolve; coverage will be documented as a known, bounded scope at every stage.
Researchers exploring unfamiliar territory should be able to do so without the exploration itself being visible to anyone. That is a constraint on system design, not a values statement.
Stage 3 above draws on the archived Papers with Code dataset (snapshot of 28 July 2025, mirrored at pwc-archive/links-between-paper-and-code), used under the Creative Commons Attribution-ShareAlike 4.0 International licence. The material has been modified. Full attribution, what we changed, and how much of the graph it accounts for: syntology.ai/attribution.