Model Kombat: what the telemetry says
A 12,000-neuron slice of the maleCNS connectome fights a
language model in mk.js. Every match writes a JSON file; this page is generated from all of them by
scripts/report_telemetry.py. Generated 18 September 2026.
The competitors
A match is always one fly against one opponent, and neither side is a single thing. On the fly's side what changes is how much the brain is allowed to change, and whether the wiring is the real one. On the other side what changes is where the decisions come from. Everything below is a record of one of these against one of those.
The fly brains
all of them the same 12,000 neurons of maleCNS, differing only in how they learn
- Wired (no learning)
- The connectome as mapped, with learning switched off. What the fly does is whatever the wiring does: senses drive motor pools, the busiest pool picks the move, nothing changes.
- In-match learning
- The same brain, with a dopamine-like rule switched on for one match. Landing a hit strengthens whatever just fired; taking one weakens it. It starts over every match.
- In-match learning + baseline
- The same, except each move is judged against its own recent average instead of against zero. Without this a losing fly suppresses everything it tries, including the moves that work.
- Curriculum
- A run of eight matches where each one starts from the best brain saved so far, so what one match learns is still there in the next.
- Curriculum + thresholds
- A curriculum that also learns how certain a motor pool must be before it is allowed to act — patience as something learned rather than set by us.
- Reloaded brain
- A brain that finished a curriculum against one opponent, entered cold against a different one. It learns nothing new on the way in; it simply arrives trained.
- Rewired control
- The null model. Same neurons, same number of connections into and out of each one, same transmitter signs, same synapse weights — only which neuron connects to which is shuffled.
The opponents
all of them answering on the same state, differing in what does the answering
- Rule bot
- Five lines of hand-written positional rules: close the distance, strike when in range, otherwise block. No model, no learning.
- Jev API
- TypeSafe's System One model, reading the fight as JSON and returning one typed move with probabilities, about five times a second.
- Local policy
- 168 weights fitted to Jev's own probabilities — the same judgment, answered in microseconds instead of a network round trip.
- Hybrid
- The local policy moving at the fly's speed while Jev corrects it live, and each correction trains the policy further.
Standings
Rounds, counted from each competitor's own point of view. Every fly has now met every opponent, so the win rates are comparable — but the rounds behind them are not evenly spread, and a fly that spent most of its rounds against the rule bot is not the same as one that spent them against the hybrid. The grid below is where that shows.
| Competitor | Rounds | W | D | L | Win rate | Damage / s | Record |
|---|---|---|---|---|---|---|---|
| Reloaded brain fly brain | 50 | 38 | 0 | 12 | 76% | 5.84 | |
| Curriculum + thresholds fly brain | 86 | 45 | 16 | 25 | 52% | 4.78 | |
| Rewired control fly brain | 53 | 25 | 11 | 17 | 47% | 2.16 | |
| Curriculum fly brain | 140 | 66 | 13 | 61 | 47% | 3.79 | |
| In-match learning fly brain | 18 | 6 | 0 | 12 | 33% | 2.89 | |
| In-match learning + baseline fly brain | 21 | 7 | 0 | 14 | 33% | 3.80 | |
| Wired (no learning) fly brain | 65 | 8 | 0 | 57 | 12% | 3.24 | |
| Hybrid opponent | 112 | 62 | 27 | 23 | 55% | 5.17 | |
| Jev API opponent | 91 | 48 | 1 | 42 | 53% | 5.06 | |
| Local policy opponent | 135 | 69 | 12 | 54 | 51% | 5.00 | |
| Rule bot opponent | 95 | 19 | 0 | 76 | 20% | 2.52 |
Who actually played whom
Every competitor against every other. The big number is the share of rounds the fly won; under it, the record as rounds won–drawn–lost. Read down a column to see what the fly's learning setup is worth against a fixed opponent; read across a row to see how far that setup carries. The tint is rounds won minus rounds lost, so a cell full of draws sits near level even when its win rate reads zero. The wiring alone loses to the whole field; a brain carried in from an earlier curriculum is the only fly with a winning record against all four.
Table: every match-up
| Fly brain | Opponent | Win rate | Won | Drawn | Lost |
|---|---|---|---|---|---|
| Wired (no learning) | Jev API | 12% | 2 | 0 | 14 |
| Wired (no learning) | Hybrid | 0% | 0 | 0 | 20 |
| Wired (no learning) | Local policy | 0% | 0 | 0 | 16 |
| Wired (no learning) | Rule bot | 46% | 6 | 0 | 7 |
| In-match learning | Jev API | 67% | 4 | 0 | 2 |
| In-match learning | Hybrid | 0% | 0 | 0 | 4 |
| In-match learning | Local policy | 50% | 2 | 0 | 2 |
| In-match learning | Rule bot | 0% | 0 | 0 | 4 |
| In-match learning + baseline | Jev API | 60% | 3 | 0 | 2 |
| In-match learning + baseline | Hybrid | 0% | 0 | 0 | 4 |
| In-match learning + baseline | Local policy | 40% | 2 | 0 | 3 |
| In-match learning + baseline | Rule bot | 29% | 2 | 0 | 5 |
| Curriculum | Jev API | 100% | 24 | 0 | 0 |
| Curriculum | Hybrid | 0% | 0 | 7 | 17 |
| Curriculum | Local policy | 35% | 26 | 6 | 43 |
| Curriculum | Rule bot | 94% | 16 | 0 | 1 |
| Curriculum + thresholds | Jev API | 20% | 4 | 0 | 16 |
| Curriculum + thresholds | Hybrid | 0% | 0 | 16 | 8 |
| Curriculum + thresholds | Local policy | 100% | 18 | 0 | 0 |
| Curriculum + thresholds | Rule bot | 96% | 23 | 0 | 1 |
| Reloaded brain | Jev API | 56% | 5 | 0 | 4 |
| Reloaded brain | Hybrid | 72% | 21 | 0 | 8 |
| Reloaded brain | Local policy | 100% | 6 | 0 | 0 |
| Reloaded brain | Rule bot | 100% | 6 | 0 | 0 |
| Rewired control | Jev API | 0% | 0 | 1 | 10 |
| Rewired control | Hybrid | 29% | 2 | 4 | 1 |
| Rewired control | Local policy | 0% | 0 | 6 | 5 |
| Rewired control | Rule bot | 96% | 23 | 0 | 1 |
The fly learns, against every opponent but one
Damage dealt per second by the fly, averaged over each generation's rounds. Each generation inherits the best brain saved so far and keeps learning. Against the hybrid — the local policy moving at the fly's own speed with Jev correcting it live — learning from scratch converges on not fighting at all.
Table: fly damage per second by generation
| Opponent | gen 1 | gen 2 | gen 3 | gen 4 | gen 5 | gen 6 | gen 7 | gen 8 |
|---|---|---|---|---|---|---|---|---|
| Rule bot | 6.04 | 8.50 | 8.38 | 8.10 | 8.53 | 7.87 | 8.46 | 8.16 |
| Jev API | 6.14 | 7.16 | 8.15 | 6.93 | 6.34 | 7.01 | 7.03 | 6.90 |
| Local policy | 1.87 | 2.90 | 2.98 | 2.95 | 2.83 | 2.93 | 4.16 | 2.93 |
| Hybrid | 1.87 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 |
Scrambling the connectome: still learns, stops landing clean hits
The control keeps every neuron, every in- and out-degree, every transmitter sign and every synapse weight, and shuffles only which neuron connects to which. It still learns to beat the rule bot — at 2.3 damage per landed hit against 5.5, which is what a blocked hit is worth.
Table: real vs rewired, damage per second
| Circuit | gen 1 | gen 2 | gen 3 | gen 4 | gen 5 | gen 6 | gen 7 | gen 8 |
|---|---|---|---|---|---|---|---|---|
| Real connectome | 6.76 | 8.00 | 7.42 | 7.55 | 7.56 | 7.55 | 7.94 | 7.97 |
| Rewired control | 3.05 | 3.80 | 4.40 | 4.16 | 3.93 | 3.90 | 4.04 | 3.84 |
What learning changes, synapse by synapse
Mean strength of the excitatory synapses onto each motor pool, relative to the wiring (1.0 = as mapped). The guard grows; the escape jump — a real circuit doing its job, useless against someone who keeps punching — is suppressed.
Table: learned strength by pool
| Pool | gen 1 | gen 2 | gen 3 | gen 4 | gen 5 | gen 6 | gen 7 | gen 8 |
|---|---|---|---|---|---|---|---|---|
| Wings (guard) | 1.40 | 1.49 | 1.42 | 1.52 | 1.55 | 1.53 | 1.57 | 1.53 |
| TTMn (escape jump) | 0.96 | 0.96 | 0.96 | 0.96 | 0.96 | 0.96 | 0.96 | 0.96 |
| Hind legs (kick) | 1.04 | 1.06 | 1.03 | 1.06 | 1.07 | 1.06 | 1.06 | 1.06 |
Every round played, by opponent
All rounds in the telemetry, including the runs where the fly was learning from scratch and lost badly.
Table: round outcomes
| Opponent | Fly won | Drawn | Fly lost |
|---|---|---|---|
| Rule bot | 53 | 0 | 18 |
| Jev API | 40 | 0 | 24 |
| Local policy | 54 | 6 | 64 |
| Hybrid | 21 | 23 | 61 |
| Rule bot · rewired | 23 | 0 | 1 |
| Jev API · rewired | 0 | 1 | 10 |
| Local policy · rewired | 0 | 6 | 5 |
| Hybrid · rewired | 2 | 4 | 1 |
What each side actually does
Share of decisions by move, from 3,645 logged decisions. Jev spreads itself across closing, striking and blocking; the fly, as wired, spends a quarter of its moves jumping away.
Table: decision mix
| Move | Jev | Fly brain |
|---|---|---|
| stand | 0% | 0% |
| walk_forward | 23% | 28% |
| walk_backward | 1% | 0% |
| jump_away | 0% | 17% |
| punch | 24% | 0% |
| kick | 26% | 20% |
| block | 26% | 34% |