Model map
Every model in the corpus, placed by how its freeflow prose resembles every other model's. Distance here reflects wording, topics and style, not capability or ancestry: the character n-gram centroid of each model's freeflow samples, compared pairwise by cosine, then projected down to something you can rotate.
Similarity tree
A hierarchical summary of the full-space distances. Merge heights reflect average-linkage distances between clusters; this tree does not preserve every pairwise distance or establish model ancestry. Leaves are coloured by lab, so you can see where the tree recovers labs on its own — and where it doesn't.
Method and honesty
Across 154 model identities, 30,820 freeflow responses are vectorised as within-word character 3–5-gram TF-IDF and averaged into one centroid per model. Each response has equal weight within its pooled model; routes and conditions with more responses therefore contribute more. The vocabulary and IDF weights are fitted across all responses. Cards show sample counts and collection cells; these are not equally precise estimates. A model label may pool multiple routes or captures, and a router may expose different backends to different prompts.
Cosine similarity between centroids gives the 154×154 reference matrix for this visualisation—not ground truth about personality or training. The map necessarily loses structure. Its retention score is the average fraction of original top-three neighbours retained in projected coordinates. Individual models can retain anywhere from none to all three; the selected card shows its own count. Scores are calculated at the precision of the shipped data, with stable model-order tie breaking. Axes use equal units; perspective can still alter apparent distances on a 3D screen.
- UMAP 2D 57% · 3D 61% — prioritises local neighbourhoods; cluster gaps are not original distances
- MDS 2D 36% · 3D 50% — approximates global distances; relative distance error 0.220 / 0.140
- PCA 2D 28% · 3D 38% — linear projection of unit-normalised centroids; explains 27% / 34% of their feature variance, not of pairwise distances
Relative distance error is the root-sum-squared pairwise residual divided by the root-sum-squared original distances, after fitting one uniform scale. Lower is better. It complements neighbour retention, rather than making MDS lossless. Read the drawn neighbour edges and card similarities as measurements in the full feature space, the tree as a hierarchical summary, and the map positions as an approximation. These are exploratory estimates; no bootstrap confidence intervals or cross-seed stability guarantees are claimed.
Proximity describes similar output distributions under these collection conditions. Shared training, post-training, distillation, prompting, vocabulary, topic choices and routing are possible explanations—not conclusions established by this plot. It does not establish developer identity, ancestry, capability or subjective experience. Lab labels only colour the result; they do not determine the embedding. Models with an unknown lab are ringed.
Release sequences connect a lab's or family's models in release-date order, not a lineage. Same-day ordering is arbitrary. Time filtering uses this single present-day embedding, including the influence of later models; it is not a reconstruction of what the map would have looked like then. Release dates are not capture dates.
Reproducibility and collection scope
Generator: website/scripts/generate_map.py. Seed 0;
UMAP neighbours 8, minimum distance 0.15.
PCA uses unit-normalised centroids. Missing UMAP dependencies stop generation rather than publish incomplete data.
Versions: numpy 2.3.5 · scikit-learn 1.8.0 · scipy 1.17.1 · umap-learn 0.5.12.
Input fingerprint: f6d7d3de4c72c558586ac3a8eed89ac09e2f66ae07206215af9c904884370ce5
Generator fingerprint: 32bec8b04cd51d7a8b6bf9a258ec87b320de2809a5fd7708a9f4e5824a605e5d
Repository base revision: 6417be8d3cb4461ad529000dafa2a08da05c0f51.
Sample/cell counts describe the responses pooled in each centroid. Capture dates are shown only when explicitly present in the published samples; missing dates are not inferred from release dates or file modification times. Data: this corpus.