The result
A recurrent controller learned to query a synthetic world through a hard one-cell read interface. On 16-cell pointer chains it reached 100% accuracy at depths two and four, and 98% at depth eight. On direct lookup, a two-phase training schedule reached 100% accuracy with exactly one mean read at both 16 and 64 cells. That is the information-theoretic minimum for a one-cell lookup.
The same model did not scale. Accuracy fell to at most 4% at 128 cells and to roughly zero at 256 cells. Depths 16 and 32 also failed. The first cycle localized two causes: address bits above the trained range never varied, and the learned policy horizon stayed tied to the ten-step training schedule.
A second experiment attacked the address-bit explanation by drawing identifiers from the full 16-bit space. It made the task harder in distribution instead of restoring extrapolation. At 16 cells, accuracy ranged from 0.3375 to 0.50 across three seeds. At 64 cells it ranged from 0.0125 to 0.0525. At 1,024 cells, nearly every grid cell was zero.
What the controller actually learned
The controller was a GRU with a learned query head. Each query selected one cell by forward hard argmax; the backward pass used a soft straight-through estimator. Inputs contained the question bits and prior reads. Tests excluded answer, chain, difficulty, and hop-count side channels.
The positive result was not a fixed schedule in disguise. Corrupting a read changed the next query in 99.5% of examples. A mid-chain pointer edit redirected trajectories in 70.5% of cases. Irrelevant growth from 64 to 256 cells changed the read count by only -0.005.
The policy therefore learned conditional information acquisition. It also learned a brittle addressing procedure bound to the training distribution. Those are separate facts.
Why the first headline was too strong
The first cycle had four limits. It used one seed, the full-access baseline had 30% more parameters than the fixed reader, the XOR tasks were too hard for every arm, and the stronger Transformer baseline was never run.
The seed follow-up introduced another confound. Seed 0 trained for 15,000 steps; seeds 2 and 3 trained for 8,000. The large seed-0 gap at 64 cells mixes seed sensitivity with training length. Only the seed-2 and seed-3 comparison isolates seed variance, and those two runs agree in their failure.
The bounded claim is clear: stochastic gradient descent can discover an adaptive policy through a hard read channel on tasks constructed to reward adaptivity. The current controller does not establish size or depth generalization, and it does not establish an advantage over a compute-matched Transformer.
What changed next
The next experiment should separate representation from optimization. Keep random identifiers, match training steps across seeds, add a Transformer reader with the same compute ledger, and report confidence intervals. If the controller still fails at trained sizes, world-keyed content addressing is the immediate bottleneck. If it fits trained sizes but fails beyond them, the extrapolation problem remains.
Low read count is not useful when the policy reads the wrong cell. A compute ledger must travel with task success.
Evidence ledger
Executed: direct lookup, pointer-chain, corruption, redirect, irrelevant-growth, identifier, and seed experiments.
Read: the primary result ledger and round-four report in
settling-field-lab.Not established: out-of-distribution scale generalization or superiority to a compute-matched Transformer.
Known confound: unequal training steps between seed 0 and seeds 2 and 3.



