Early ECP calibration experiments
This route preserves the historical place of the earlier pilot studies. Those experiments helped shape later controls, but they are not the strongest current evidence for the ECP architecture.
What the pilot contributed
The early studies suggested that a short calibration exchange could alter measured model behavior in routing tasks. They also exposed the need for stronger controls separating prompt length, multi-turn interaction, unusual tokens, and the complete contradiction-holding architecture.
Why it is not the lead claim
Small samples, task-specific outcomes, and less complete causal isolation limit the conclusions that should be drawn from the pilot. It should be read as developmental evidence, not proof of universal calibration effects or a general complexity law.
The current standard is the four-arm study.
Two hundred independent runs compare baseline, neutral multi-turn, token-only, and full CONEXUS conditions with explicit statistical limits.
Review the Current Causal Evidence