Evidence-gated execution
Reliable automation increasingly depends on deterministic controls outside the model. A setup-security study found that agents usually missed malicious package sources across ecosystems; prompts only helped for the attack dimension they named, while a pre-install check covering names, sources, and versions closed most of the observed gap. Proof-or-Stop applies the same principle to lifecycle decisions: claims such as “tested” or “done” advance only with fresh evidence bound to the tracked source state. Its gated loop recorded zero false-done outcomes in 10 scenarios and rejected all 18 tested receipt-tampering classes, though evaluation was limited to one model family and a self-hosted corpus.