Corpus expertise evaluation
Machine Studying gives a concrete test for an agent’s ability to prepare on an unseen document corpus before downstream tasks are revealed. The proposed StudyBench covers DSPy code, OpenClaw code, and machine-learning literature. Its metric rewards accuracy at lower inference-token budgets, so an agent that needs many search loops receives less credit.
The early result is cautious. Retrieval-augmented generation, long context, and simple fine-tuning do not reliably create usable corpus expertise. One example shows Qwen3.5-9B improving on DSPy when forced to use 20 search iterations, but the broader point is that accessible evidence can remain unused without better study behavior.