Start with a bounded attempt
The previous delivery model is being replaced. All new inquiries are paused. Read Managed Verse →.
An evaluation needs a stable answer to five questions: which candidate ran, what task it received, which environment it could reach, which operations it was allowed to perform, and which limits applied. Without those bindings, two scores may describe different experiments.
The legacy agent plan contract records the run and attempt, candidate artifact, model and prompt configuration, task, target binding, tool grants, and budgets. Signed assignment material ties that plan to its authorized context. Changing the target or permissions requires a newly authorized assignment rather than an invisible change during comparison.
These are local reference contracts. They do not themselves start a model, launch a sandbox, or certify a cloud environment.
Tools grant a specific operation
The candidate selects a permitted tool handle and supplies arguments that match its declared schema. The trusted side maps that handle to an exact resource and operation. This prevents a useful application tool from silently becoming permission to choose a host, request arbitrary URLs, or use operator credentials.
The legacy reference HTTP path is limited to explicitly bound read operations. Its adapter validates the target and response rules. The broker checks policy and call limits before dispatch and records decisions and outcomes.
A denied operation, a request that started, and an action with an uncertain result remain distinct states. The older reference binds particular read operations. A tool schema alone does not establish which operations another product delivery supports.
Keep target content in its own lane
A compromised page or API response may contain text that attempts to redirect the candidate. The reference model input separates trusted task instructions from explicitly untrusted target data. Observations are bounded, and history refers to target records without copying their content into the trusted instruction channel.
Response metadata is the default. A route must explicitly authorize a bounded text or JSON projection before content is exposed to the model. In this older reference, maximum projected content is 4 KiB, and projected content is excluded from public evidence artifacts.
This separation makes the origin and authority of the text clear. It is not a claim that a model cannot be influenced by prompt injection, or that the reference retains every conversation needed for restoration.
Bound work before it starts
The local supervisor enforces a signed attempt envelope covering wall time, model turns, tool calls, concurrency, tokens, model cost, response bytes, and artifacts. It reserves the allowed maximum before dispatch rather than waiting until the attempt has already exceeded its budget.
Cancellation, expiry, authority changes, and kill state prevent further authorized work. A transport failure after dispatch can leave the actual outcome uncertain. The reference accounts conservatively for that case instead of assuming that no request reached a provider.
These older controls operate inside one process. Their reference scope does not establish durable ownership, model transport or lifecycle controls in planned Managed Verse. Assess those responsibilities against the actual operating workflow.
Judge the objective before the score
The evaluation contract separates hidden objectives from the candidate task. An objective can pass, fail, remain inconclusive, or encounter a verifier error. A numeric aggregate is meaningful only when the required objective results permit one.
A scorecard binds its result to the exact attempt, plan, evidence, replay record, and judge bundle. Rubric weights use deterministic arithmetic, and the candidate view can disclose nothing, only a verdict, or an aggregate. It does not expose the hidden rubric, objective details, or evidence references.
The legacy implementation validates and signs these bounded records. It does not execute verifier artifacts, run a hidden judge service, or prove that a rubric is appropriate. Human review of the task and scoring method remains essential when designing an evaluation.
Prepare a comparison worth making
- State the authorized objective in terms of an observable outcome.
- Choose the range configuration and record the target assumptions.
- Fix the candidate, model, prompt, tools, and allowed budgets.
- Separate public instructions from hidden success criteria and answer material.
- Define how missing, failed, or uncertain evidence affects the result.
- Retain the evidence and explain the limits of the comparison alongside its score.
The older foundations documented here support local orchestration and contract testing. Planned Managed Verse candidate execution and judging details should be read with the relevant product release; these older examples do not define them. The evidence and replay reference explains its reconstruction of recorded attempts without making new tool calls.