Skip to content
andydataguy
A corpus, and what reads it

Every number on this page was measured, and every part of the system says which state it is in.

This is a walk through two knowledge environments: how much text is in them, what shape that text takes, what the retrieval layer does with it today, and where the work stops. 8 components run against the real corpus. 2 run with limits that are named. 6 are code with no data behind them yet, and they are labelled that way everywhere they appear.

The corpus and what reads itvectorfull text3 hits, sourced127 of 261 bars clear the nullsummary tierone altitude04 / the gap
Four beats. The corpus splits into chunks on its own headings. A query runs two retrieval legs, a vector leg and a full-text leg, and fuses them into one ranked answer where every hit carries the document it came from. A topology probe finds real cyclic structure in the embedding cloud and scores it against 10 shuffled nulls. Above the chunk there is a summary tier, drawn dashed because it is designed and coded and holds no data.
Substrate figures measured 2026-07-26, topology 2026-07-17.
What is built

Sixteen components, and the state each one is actually in

A retrieval stack is easy to describe and hard to verify, so this page starts with the ledger. Every capability named further down carries its state from this table, and the page fails to build if a section describes a component the ledger does not list.

running8
Executes against the real corpus on the live deployment today.
runs with limits2
The code is real and runs, and something about its coverage or its inputs keeps it from applying to every query.
code only, no data6
The code exists and is bounded and tested. No data has ever flowed through it.
Ingestion, chunking, embedding, persistence
running

Assets land, split on markdown headings at a 1800 token target with a 2600 token ceiling, embed, and persist to the live deployment.

Embedding store
running

gemini-embedding-2-preview at 3072 dimensions on a Convex vector index, asymmetric: documents stored under RETRIEVAL_DOCUMENT, queries embedded at read under RETRIEVAL_QUERY.

Governance filter ahead of ranking
running

Sensitivity is enforced per chunk before any text reaches the ranker, and the filter fails closed.

Hybrid retrieval with rank fusion
running

A vector leg and a full-text leg run against the same corpus and fuse by reciprocal rank, rather than one leg deciding alone.

Provenance on every hit
running

Each returned chunk carries its asset id, chunk id, title, and the document it resolves to.

Entity graph over the corpus
running

Typed entities extracted from the real corpus, stored as nodes and edges in Convex and rendered in the portal.

Corpus topology probe
running

A persistent-homology run over corpus embeddings, scored against a shuffled null rather than asserted.

Connector and gap detection
running

Betweenness scoring for connectors and inter-community gap detection share the community module. The gap table holds 1,534 rows against the live graph.

Community detection over the entity graph
runs with limits

Louvain runs on a fixed seed and its 281 communities across 22 groups are fully enumerated, each with size, cohesion, member names, and label terms. Three groups carry partitions describing graphs that no longer exist and nothing marks them stale, so a consumer has to reconcile against live node counts before trusting a group.

Graph-aware re-ranking
runs with limits

Personalized PageRank over the entity graph is wired into the fusion, and it returns empty and drops out whenever the graph is not graduated for that group. It contributes to some queries and not others.

Hierarchical summary tree, builder
code only, no data

No clustering implementation exists. The dimensionality-reduction and mixture-model settings appear as configuration field names with defaults and nothing reads them.

Hierarchical summary tree, reader
code only, no data

Bounded, vectorized retrieval code that reads a tree nothing writes. Its own docstring records that no caller constructs or passes a tree.

Hierarchical summary tree, stored data
code only, no data

No stored chunk carries a level, a parent, a child, or a leaf flag. The write path never sends those fields and the database validator does not declare them, so a correct builder would today have nowhere to put its output.

Multi-altitude query routing
code only, no data

Only the leaf level is ever queried, because only the leaf level exists.

Flat-twin comparison harness
code only, no data

The controlled comparison that would decide whether this retrieval beats flat top-k has never been run.

Portal search interface
code only, no data

Retrieval is reachable over the API. There is no search route in the portal.

On reading this table
Code existing and code having run against real data are different claims. Three of the components below have passing tests and no data, which is the most misleading green signal in the repository, so this page separates the two everywhere. A hollow diamond marks a source that proves something was written. A filled diamond marks a source that proves something ran.Measurement source: docs/recon/metagraph-map-20260726/MAP.md

The corpus, counted

Six and a half million words, and most of them were never published

Two knowledge environments, counted document by document with the stripping rules written down. The gap between what went in and what came out is the part that describes how the work is actually done.

6.65Mwords
Across 788 documents in both environments
Corpus census, 2026-07-26
886,208words
Published and reader-facing, across 140 documents
Corpus census, 2026-07-26
5.77Mwords
Source material behind the published work, across 648 documents
Corpus census, 2026-07-26
6.51:1
Source words behind every published word
Derived from the two totals above

Where the words are

SurfaceDocsWordsMedian
Looikos decks (brand + foundation)published48577,50713,757
Wiki articlespublished35211,9322,485
Essay library (published HTML)published1463,4083,726
Field notespublished1223,460471
Case studiespublished319,901329
RAW knowledgebasesource4605,119,2235,517
original_notes (personal notes and transcripts)source25231,6413,892
Project docs (briefs, recon, handoffs)source68153,269991
Reference material (source of truth)source13100,6335,922
Published surfaces in full, then the four largest source surfaces. Source sha256 d7f5e32b52c2dcd9, census generated 2026-07-26 at repo commit a6769d3d.

Six and a half source words per published word

5.77M words of notes, transcripts, knowledge base and working documents stand behind 886,208 words of finished writing. That ratio is the reason a retrieval layer exists here at all. The published surface is small enough to read; the material it was drawn from is not, and it is where an answer to a specific question actually lives.

What this census did not count

The census counted words in documents. It did not count any of the following, so no number on this page derived from it describes them:

  • chunks
  • embeddings
  • vector index size
  • retrieval neighbourhood or cluster graph
  • site chrome copy
  • binary asset byte totals

Chunks and embeddings appear elsewhere on this page. They come from a separate read of the live deployment on a different date, never from this census, and the two instruments measure different systems.

The shape

A hundred and tenfold spread, inside one site

The shortest published document is a case study at 185 words. The longest is a wiki article at 20,347. Both are finished, reader-facing work on the same property, and a retrieval layer has to serve them with the same machinery.

110x
Between the shortest published document and the longest
Derived from the per-surface minimum and maximum
185words
The shortest, in Case studies
Corpus census, per-document measurement
20,347words
The longest, in Wiki articles
Corpus census, per-document measurement
5
Published surfaces, each with its own length distribution
Corpus census
How long a document is, by surfacedecksessayswikinotescases185 to 20,347 wordswords per document, log scale: 100 / 1k / 10k
Each bar runs from a surface’s shortest document to its longest. The solid block is the middle half, the amber tick is the median, and the axis is logarithmic because the range covers two and a half orders of magnitude. The medians sit far apart while the ranges overlap heavily, and that combination is the finding: a surface has a characteristic length, and knowing it tells you little about any individual document on it.
Every value measured per document by the corpus census.

The wiki is bimodal, and the average hides it

Wiki articles average 6,055 words and the median is 2,485. That gap is the whole story: a quarter of them sit at or below 2,137 words while the top quarter starts at 11,044 and runs to 20,347. There is no typical wiki article. There are short reference entries and there are long essays, and the mean describes neither.

Why this breaks a fixed chunking assumption

A 329-word case study is smaller than a single chunk at the 1,800 token target, so it is one chunk and its embedding represents the entire document. A 13,757-word deck becomes dozens, and no single one of them represents the document at all. The same query therefore competes whole documents against fragments of documents in one ranking. That is the specific problem a summary tier above the chunk is meant to solve, and it is why the absence of that tier is the most consequential gap on this page rather than a nice-to-have.

The substrate

Text becomes chunks, chunks become vectors, and both keep their address

Retrieval quality is mostly decided before any query runs. These are the parameters that decide it here.

runningAssets land, split on markdown headings at a 1800 token target with a 2600 token ceiling, embed, and persist to the live deployment.
8,264
Chunks in the store, the unit retrieval actually ranks
Full database export, 2026-07-26
7,978
Chunks carrying an embedding
Full database export, 2026-07-26
3,072dims
Vector width, from Gemini Embedding 2
apps/api/src/wdc/embedding.py
1,800tokens
Chunk target, with a 2,600 token ceiling
provenance/nf46/run-2.json

Split on the document’s own headings

The chunker is markdown-heading-v1. It cuts at heading boundaries rather than at a fixed character count, so a chunk is a section the author wrote instead of an arbitrary window that starts mid-sentence. Each chunk keeps the heading breadcrumb it was cut under.Code source: wikidesignco/provenance/nf46/run-2.json

Documents and queries embed differently

Stored text embeds under RETRIEVAL_DOCUMENT and an incoming question embeds under RETRIEVAL_QUERY. A question and the passage that answers it are different kinds of writing, and asking one model to place both in the same space with the same instruction costs recall.Code source: wikidesignco/apps/api/src/wdc/embedding.py:19-20

A wrong-width vector raises

The embedding path checks the returned dimensionality and raises when it disagrees. A silent truncation or a quiet fallback to a smaller model degrades every downstream ranking while every dashboard stays green, which is the failure mode that takes months to find.Code source: wikidesignco/apps/api/src/wdc/embedding.py:19-20

The store is Convex, holding the documents, the vector index, and the graph tables together. There is no separate vector database and no separate graph database, so a hit and its provenance and its neighbours come out of one place.running

What a query does today

Four stages that run, and one that runs sometimes

This is the whole live path, in order. It has been exercised end to end against the real corpus on the deployed service, which is a different and much weaker claim than saying it is good.

runningLast returned a real answer on the real corpus 2026-07-16.Measurement source: wikidesignco/docs/recon/reground-20260716/PLATFORM-LIVE-LOG.md:396-416

That date is the claim, rather than today. On 2026-07-26 the endpoint answered in 22.6 seconds with a 401: the service is deployed, reachable, and authenticating, and the re-verification pass did not hold an API key. So the path is known to be up and is not known to be returning results as of today.

  1. 01

    Sensitivity is enforced before ranking

    Each chunk carries a sensitivity classification, and the filter applies at the database before any text reaches the ranker. It fails closed, so a chunk whose classification cannot be established does not get ranked and then filtered out of the response. It never enters the ranking at all.Code source: wikidesignco/apps/platform/convex/schema.ts:1238-1253

  2. 02

    Two legs run, and neither one decides alone

    A vector leg finds passages that mean the same thing as the question. A full-text leg finds passages that use the same words. They disagree often, and the disagreement is useful: an exact identifier, a product name, or a rare token is where the lexical leg wins, and a paraphrase is where the vector leg wins.Code source: wikidesignco/apps/api/src/wdc/service.py:1030-1055

  3. 03

    The two rankings fuse by rank, not by score

    Reciprocal rank fusion combines the two lists by the position a result took in each, rather than by the raw similarity numbers. Cosine similarity and a lexical relevance score are not on the same scale and adding them is arithmetic on incompatible units, which quietly lets whichever leg happens to produce larger numbers win.Code source: wikidesignco/apps/api/src/wdc/service.py:1030-1055

  4. 04

    Every hit carries where it came from

    A result returns with its asset, its chunk, its title, and the document it resolves to. That is what makes an answer checkable rather than plausible, and it is the difference between a system a buyer can audit and one they have to trust.Code source: wikidesignco/provenance/nf46/run-2.json (searchResults[].sourceRef)

05. The graph re-rank, when it applies

runs with limits

A personalized PageRank pass over the entity graph can lift results that sit near the entities a question is about. It is wired into the fusion and it returns empty whenever the graph is not graduated for that group, at which point it drops out and the other legs decide. So it contributes to some queries and not to others, and a reader should assume the fused result is the two text legs unless the graph happens to be ready.Code source: wikidesignco/apps/api/src/wdc/service.py:1069-1110

What this section does not claim

That any of the above beats a standard retrieval stack. The comparison that would establish it has not been run. The recon lane that exercised this path against the live deployment recorded its own verdict plainly: what it had just run is flat top-k retrieval, and nothing it saw entitles anyone to call it better than standard. The governance filter, the fusion, and the provenance are real and they are more than an off-the-shelf assembly gives you. Being more than standard in construction is not the same as being better in results, and only one of those has been measured.Measurement source: wikidesignco/docs/recon/reground-20260716/PLATFORM-LIVE-LOG.md:418-424

Neighbourhoods

The system grouped the corpus on its own, and gave the groups names

Alongside the vectors there is a graph of the things the corpus talks about. The neighbourhoods in it are computed rather than drawn by hand, and each one carries label terms derived from its members.

3,112
Entity nodes across the graph
Live export, 2026-07-26
16,013
Edges between them
Live export, 2026-07-26
227
Communities across the 19 groups whose partition still matches their live graph
Live export, 2026-07-26
2,121
Extracted statements, beside 1,534 detected gaps and 17,028 enrichments
Live export, 2026-07-26

Named, not numbered

Every community carries label terms taken from its members. These three are real, rendered as stored. Nobody supplied a taxonomy or defined these groupings. They fall out of which entities occur together across chunks.

  • solo quantstuck fundsprawl traderDCA investorburned user

    A customer-segment cluster. Nobody defined these segments or asked for them. They fell out of which entities co-occur.

  • WikiDesignCoSuperHarnessSymphony AGIAndyHermes

    The named systems and the person they belong to, grouped together.

  • Wardley axisAnimation 6dAI generation tools

    A method-and-tooling cluster spanning several documents.

One group, fully enumerated

runs with limits

wikidesignco_kb partitions 497 nodes and 3,429 edges into 15 communities. The largest holds 117 members, the median 22, the smallest 4. Its member count reconciles exactly against the live node count.Code source: wikidesignco/apps/api/src/wdc/retrieval/analytics_build.py:204-209

networkx Louvain, resolution 1, seed 42, threshold 1e-07, with modularity stored beside the run rather than asserted (0.735 on node-foreman_kb). The same parameters reproduce the same partition.

“Seeded Louvain is reproducible for this graph and parameter set, but the partition is a heuristic rather than an ontological fact.”

The system files its own result as heuristic and says so in the record. The partition is flat: one level of neighbourhoods, with nothing nested above it.Measurement source: wikidesignco/docs/recon/reground-20260716/STITCH-MAP.md:408

Real entities, typed

running

Entities are extracted from the corpus and stored as nodes and edges beside the chunks they came from, in the same database. Folding case, the three largest types are Concept 821, Framework 442, Organization 426.Measurement source: wikidesignco/docs/recon/reground-20260716/PLATFORM-LIVE-LOG.md:378-380

The extraction pass normalizes entity names and does not normalize the type field, so each type is stored under two spellings. These totals fold case. An unfolded count reports roughly half of each.

54 communities excluded from the count above

Three groups hold partitions describing graphs that no longer exist, and nothing in the database marks them stale. One partitions 550 members against 122 live nodes, and two more hold communities against a graph with zero nodes. A surface that read the community table without cross-checking live node counts would show neighbourhoods for graphs that are gone, so this page reconciles every group first and renders only the 19 that match. The excluded ones are counted here rather than dropped quietly.

The one measured result

The corpus has real shape, and the test could have said otherwise

Anyone can assert that their corpus has structure. The only way that claim means anything is to state in advance what a corpus with no structure would look like, and then check. That is what this probe did.

The real run against ten shuffled nulls10 nullsreal0longest H1 persistencez = 102.4
Persistence of the longest one-dimensional loop. The ten grey bars are ten runs on shuffled data. The violet bar is the real corpus. The real run does not merely sit far from the null average, it sits outside the null range entirely: every one of the ten shuffles topped out between 0.0084 and 0.0125, and the real corpus reached 0.1251.
ripser over 318 distinct vectors, cosine distance, maximum dimension 1. Run 2026-07-17.

What the null destroys, and what it keeps

Each of the 318 vectors has 3,072 coordinates. The null shuffles every coordinate independently across the points. That destroys every relationship between one document and another while leaving each dimension’s own distribution exactly as it was. So whatever survives in the shuffled cloud is what this geometry produces with the meaning stripped out and nothing else changed, which is the comparison that makes the real number mean something.Measurement source: wikidesignco/docs/recon/reground-20260716/BARCODE-PROBE.md

The correction that makes the number smaller

The shuffled cloud is tighter than the real one, with a diameter of 0.42 against 0.5583. A bigger cloud can hold a longer loop for reasons that have nothing to do with meaning, so the raw 12x separation flatters the result. Normalized by diameter it is 0.224 against 0.024, which is 9x. That is the number worth quoting.

318
Distinct vectors, after collapsing duplicate chunk text
BARCODE-PROBE.md, 2026-07-17
261
One-dimensional loops detected in the real cloud
BARCODE-PROBE.md, 2026-07-17
127
Loops longer than the longest any of the ten nulls produced
BARCODE-PROBE.md, 2026-07-17
9x
Separation after correcting for cloud diameter
BARCODE-PROBE.md, 2026-07-17

The first reading was wrong

Looking at the barcode by eye, the lane called it a blob. The longest bar was only about six percent longer than the second, and the shape did not match what the usual visual heuristic wants. The written note was that the reading was an eyeball, and that “long” means nothing without a null. Running the null reversed the conclusion. There is no single dominant loop, and there is a band of roughly a dozen strong ones plus a long tail, sitting far outside anything noise produces.

The loop resolves to four documents

A barcode nobody reads is a plot rather than a finding, so the points carrying the longest loop were pulled and opened. They are 5 points across 4 different documents, sitting at almost the same section heading in each: Growth Vectors. The loop is a conceptual ring that genuinely runs across the corpus, and the documents forming it can be named and quoted.

  • 03-looikos.md
  • 12-glass-porcupine.md
  • 11-dyson-forge.md
  • 04-grid-trade-pro-deep-dive.md

One caveat that survives into the claim: ripser returns cohomology representatives, so those points localize the class rather than enumerate a ring in traversal order. They name the neighbourhood the loop lives in.

The frontier

There is one altitude, and the tier above it is not built

The design calls for summaries above the chunk, so a broad question can be answered at a level of abstraction that matches it. The reader for that structure is written. Nothing has ever written the structure.

0
level
of 8,264 chunks carry which altitude a chunk sits at
0
parentIds
of 8,264 chunks carry the summary above it
0
childrenIds
of 8,264 chunks carry what it summarizes
0
isLeaf
of 8,264 chunks carry whether it is a leaf

The reader exists

code only, no data

It is bounded, vectorized retrieval code that declares 4 altitudes and can reach none of them. Its own docstring records that no caller constructs or passes a tree to it.Code source: wikidesignco/apps/api/src/wdc/retrieval/raptor_retrieval.py:24

The builder does not

code only, no data

There is no clustering implementation anywhere. The dimensionality reduction and mixture model that the design names appear as configuration field names with default values, and nothing reads them.Code source: wikidesignco/apps/api/src/wdc/retrieval/raptor_models.py:11-13

Two independent breaks

code only, no data
  1. 1. The write path sends seven fields per chunk and none of them is a parent, a child, a level, or a leaf flag.
  2. 2. The database validator does not declare those fields at all, so a correct builder would today have nowhere to put its output.

Code source: wikidesignco/apps/api/src/wdc/ingestion/persist.py:109-120Code source: wikidesignco/apps/platform/convex/ingestion.ts:390-405

The tests pass, and they prove nothing

The property tests for this component are green. They build synthetic nodes in memory and exercise the reader against those, so they establish that the code is internally correct and say nothing at all about whether any real chunk has ever carried a level. Green tests on a component with no data are the most persuasive wrong signal in the repository, which is why this page separates code sources from measurement sources everywhere.Measurement source: wikidesignco/docs/recon/reground-20260716/LINEAGE-GATE.md

What does exist, and is easy to mistake for it

Each chunk carries the heading breadcrumb it was split under, which places it inside its own document. A breadcrumb orders one document. It does not relate two documents, and no summary sits above it. It is named almost exactly like the thing it is not, which is worth knowing before reading the schema.Code source: wikidesignco/apps/api/src/wdc/ingestion/extract.py:103-119

The four questions that motivate building it

These are the query classes flat top-k retrieval is expected to struggle with, and the reason the summary tier is on the roadmap at all. Every one of them is an argument from structure. None of them has been demonstrated here, because three of the four need a representation above the chunk and there is no representation above the chunk.

Thematic
What are the recurring themes across everything here?

No single chunk contains a theme. A theme is a property of a set, and answering it needs a representation of the set.

Multi-hop
How does the thing in document A bear on the decision in document C?

The chain runs through a document that mentions neither endpoint, so a ranking of passages against the question can miss the middle entirely.

Holistic
What is this corpus actually about?

Top-k returns the k passages most like the question. A question about the whole has no most-similar passage.

Negative space
What is missing that should be here?

Retrieval returns what exists. Absence has no embedding, so finding it needs a model of what the shape should have been.

Zero of four demonstrated

The comparison that would settle it is a flat twin: the same model, the same persona, the same corpus, two retrieval paths, scored on these four classes. It has never been run.Measurement source: wikidesignco/docs/recon/reground-20260716/PLAYTEST-CHARTER.md and REALIGNMENT.md

code only, no dataMulti-altitude routingcode only, no dataFlat-twin harnessLineage counts read 2026-07-26.
What the census found

Most of what was ingested never became retrievable, and the dashboard was right the whole time

These are defects in the operator's own corpus. They are on this page because they are the strongest evidence for the argument the rest of the page is making.

433
Assets ingested
Full database export, 2026-07-26
189
Of those, assets that produced no chunks and so hold no retrievable text
Full database export, 2026-07-26
244
The working corpus: assets that actually reached the retrieval layer
Full database export, 2026-07-26
8copies
Times 03-looikos.md was ingested, as separate assets of 30 chunks each
Duplicate detection, 2026-07-17

An asset count is not a corpus size

Quoting 433 documents would overstate the retrievable territory by roughly four times. The status distribution cross-checks it exactly, with no drift:

indexed
233
uploaded
188
failed
10
enriched
2

Measurement source: wikidesignco/docs/recon/reground-20260716/LINEAGE-GATE.md

Duplication, scoped to where it was measured

In the wikidesignco workspace, 545 embedded chunks collapse to 318 distinct texts, leaving 227 duplicate vectors. Copies were confirmed at euclidean distance 0 between copies, so this is exact repetition rather than near-similarity. This figure is that one workspace and has not been measured across the rest.Measurement source: wikidesignco/docs/recon/reground-20260716/BARCODE-PROBE.md

Why this is the argument, not an admission

Both defects were invisible to the platform’s own row counter, and the row counter was correct: it matched real row counts across all twenty-six workspaces with zero drift. It is an accurate counter of the wrong thing. Duplication inflates row counts, and a missing hierarchy leaves them untouched, so neither defect could ever have surfaced there. A purpose-built instrument found both in one pass. That is the case for building the instrument instead of trusting the dashboard, and it is the same case every other section here is making.

Every source

Nothing here rests on this page's word

Each claim above carries a marker linking to its entry below. A hollow diamond marks a source that establishes something was written. A filled diamond marks a source that establishes something ran against real data. Only the second kind supports a claim that a component works.

The corpus, counted
  • wikidesignco/docs/recon/reground-20260716/LINEAGE-GATE.md

    The live corpus counts, and the count of chunks carrying tree lineage fields.

Ingestion and embeddings
  • wikidesignco/provenance/nf46/run-2.json

    The chunker name and its token target and ceiling.

  • wikidesignco/apps/api/src/wdc/embedding.py:19-20,36-38, guard at :94-98

    The embedding model, its dimensionality, the asymmetric task types, and the dimension guard that raises rather than degrading quietly.

  • docs/recon/metagraph-map-20260726/MAP.md

    The build-state ledger this page renders.

Retrieval
  • wikidesignco/apps/platform/convex/schema.ts:1238-1253, enforcement stated at apps/api/src/wdc/retrieval/__init__.py:8-10

    Sensitivity enforced per chunk before ranking, failing closed.

  • wikidesignco/apps/api/src/wdc/service.py:1030-1055, retrieval/fusion.py

    The vector leg and the full-text leg fusing by reciprocal rank.

  • wikidesignco/provenance/nf46/run-2.json (searchResults[].sourceRef)

    The asset, chunk, title, and resolved document carried per hit.

  • wikidesignco/docs/recon/reground-20260716/PLATFORM-LIVE-LOG.md:396-416

    End-to-end retrieval running on the deployed service against the real corpus.

  • wikidesignco/apps/api/src/wdc/service.py:1069-1110

    Graph-aware re-ranking, and the condition under which it returns empty and drops out of the fusion.

  • wikidesignco/docs/recon/reground-20260716/PLATFORM-LIVE-LOG.md:418-424

    The recon lane's own verdict on what the live retrieval currently is.

The entity graph
  • wikidesignco/docs/recon/reground-20260716/PLATFORM-LIVE-LOG.md:378-380

    The entity graph populated, and the rendered connection count.

  • wikidesignco/apps/platform/convex/schema.ts:1613-1660, 1629-1639

    The node and edge tables, and that the top-N read is a caller-supplied level-of-detail cap rather than the population.

  • wikidesignco/apps/api/src/wdc/retrieval/analytics_build.py:204-209, betweenness at :328, determinism at :45-46

    Community detection and connector scoring, computed on a fixed seed rather than hand-authored.

  • wikidesignco/docs/recon/reground-20260716/STITCH-MAP.md:408

    That the community partition is flat rather than nested.

The topology probe
  • wikidesignco/docs/recon/reground-20260716/BARCODE-PROBE.md, with barcode_probe.py and null_test.py beside it

    The topology probe: the vector set, the metric, the bar counts, the ten-trial null, the z-score, and the documents the longest loop resolves to.

The hierarchy that is not built
  • wikidesignco/apps/api/src/wdc/retrieval/raptor_retrieval.py:24,99,129,159,183

    The four altitudes the reader declares, and its own record that no caller constructs or passes a tree.

  • wikidesignco/apps/api/src/wdc/retrieval/raptor_models.py:11-13,35-50

    That the clustering settings exist as configuration field names with no implementation reading them.

  • wikidesignco/apps/api/src/wdc/ingestion/persist.py:109-120

    The seven fields the write path sends per chunk.

  • wikidesignco/apps/platform/convex/ingestion.ts:390-405

    That the database validator does not declare the lineage fields, so a builder would have nowhere to write them.

  • wikidesignco/apps/api/src/wdc/ingestion/extract.py:103-119

    That the stored hierarchy path is a heading breadcrumb within one document.

  • wikidesignco/docs/recon/reground-20260716/PLAYTEST-CHARTER.md and REALIGNMENT.md

    The definition of the four query classes, and that the comparison against flat retrieval is unstarted.