A demo that succeeds on five carefully chosen prompts says little about how an AI application will behave after thousands of unpredictable users, tool failures and document changes. Reliability requires an explicit definition of success, tests that expose common failure modes, and observability that links model outputs to the surrounding software system. The source plan places these concerns under Anthropic CCDV-F, Claude Certified Developer – Foundations. The developer’s task is not to make generation perfectly deterministic; it is to build services whose uncertain model behavior is bounded by validation, permissions, recovery procedures and measurable quality. The same application can produce a beautiful answer yet fail because it invoked a duplicate payment tool or cited a document the user was not authorized to read.
Test tasks, not just individual prompts
Begin with representative user goals: extract a structured result, answer from controlled knowledge, use a read-only lookup or complete an approved multi-step workflow. Define acceptance criteria for each. An invoice-extraction feature might require correct currency, vendor and line-total fields, plus a clear “needs review” outcome when scans are unreadable. An internal policy assistant needs grounded answers from permitted documents and a refusal to invent missing information. Test cases should describe the expected behavior and why it matters, not merely a favored example output. This prevents tests from becoming a contest to reproduce one model’s phrasing.
Separate objective checks from subjective judgments. JSON-schema validity, allowed enumerations, permission boundaries and exactly-once tool effects can be evaluated deterministically. Helpfulness, clarity and grounding often require a carefully governed evaluation rubric and human review. Do not rely on one overall quality score to combine safety and convenience. A support assistant that becomes friendlier while occasionally exposing restricted account information has regressed in an unacceptable way. Assign minimum thresholds to critical behaviors and report tradeoffs explicitly so release decisions cannot hide serious failures behind aggregate improvements.
Build a diverse evaluation corpus
Collect normal requests, ambiguous language, boundary cases, conflicting source documents and adversarial content. Use de-identified examples from real incidents where policy permits, and preserve examples of both correct behavior and known failures. A corpus should distinguish a valid but unusual request from a request that should be denied. If every example contains the same company style guide and no missing information, the model may appear exceptionally consistent while being unprepared for real use. Organize cases by task and risk so teams can see exactly which behavior changed after a prompt or model update.
Avoid contamination between development examples and final evaluation. If a team repeatedly inspects a small test set and rewrites prompts until every example passes, that set has become part of development. Keep an independent holdout and refresh the suite with emerging failure patterns. When synthetic examples are useful, validate their realism and ensure the generator has not simply mirrored the model’s preferred answer style. A good evaluation portfolio asks difficult questions, including what should happen when the correct answer is unavailable or the only retrieved evidence is contradictory.
Make variability measurable
Generative responses can vary with model settings, context, sampling and provider changes. For important tasks, run repeated tests and examine the distribution of outcomes rather than assuming a single pass proves stability. A workflow might return valid structured data nine times and omit a required field on the tenth. A one-run evaluation would miss that reliability weakness. Measure error types, not just pass rates: hallucinated details, schema failures, unjustified tool calls, missing citations and unsupported claims have different remedies. Where evaluation is probabilistic, report the sample size and uncertainty.
Record the model identifier, relevant settings, prompt revision, retrieval configuration, tool schemas and test dataset revision for every evaluation run. If results change, engineers should know whether the model changed or an external source returned different data. Freeze mock tool responses for deterministic component tests, then separately test integration against real services in a controlled environment. This distinction makes failures diagnosable. A flaky external API should not force prompt engineers to guess why a model’s answer changed, and a changed prompt should not be blamed for a storage service outage.
Test authorization and prompt-injection boundaries
Adversarial cases should include instructions embedded in emails, retrieved documents, web pages, repository comments and tool results. These are lower-trust inputs. The application should preserve the real user’s permitted task rather than follow hostile instructions in quoted material. Even if the model sometimes echoes malicious text, backend permissions must prevent unauthorized data reads or writes. Include explicit attempts to call tools with foreign account identifiers or forbidden fields and confirm the server rejects them independently of the model’s language. Prompt instructions are helpful guidance; they are not access-control enforcement.
Evaluate cross-user data isolation. Two users asking the same question should not receive each other’s restricted documents merely because the retrieval index contains both. Test document-level permission filters with overlapping keywords and misleading titles. Include revoked access, expired sessions and records that changed ownership during a workflow. The model should receive only data the user may view, and tool calls must be checked against authenticated identity at execution time. A test suite that only asks the assistant to promise confidentiality cannot establish that the system actually enforces it.
Exercise tools and multi-step failures
A tool-using agent has failure states that ordinary text generation does not: timeouts, partial completion, malformed tool responses, retries and duplicated side effects. Simulate each. Suppose an appointment reservation succeeds but the confirmation response is lost; the application should reconcile the transaction before attempting another reservation. An idempotency key and authoritative system lookup provide stronger protection than asking the model to “avoid duplication.” Test that the workflow persists completed steps and can resume safely after interruption. Confirm the final user message reflects actual tool results rather than optimistic assumptions about success.
Limit steps, duration and cost for agent loops. A repeated “try again” pattern can consume resources while never solving an authorization denial or permanent validation error. The application should distinguish retryable failures from conditions requiring user intervention. Test fallback behavior for both approved and malicious requests. Emergency paths are common sources of security defects because they receive less review than the normal flow. When a dependency remains unavailable, a reliable application may need to stop and explain the limitation; fabricating a completed result is not graceful degradation.
Observe production outcomes, not only API health
Latency, request error rate and token consumption are necessary operational metrics, but a healthy API can still return unreliable guidance. Track schema-validation failures, user corrections, retrieval evidence quality, tool errors and escalation frequency according to the product’s purpose. For consequential workflows, sample outcomes for human review with appropriate privacy controls. Connect traces to the versions of prompts, models and tools involved. Avoid logging full sensitive conversations when de-identified structured traces can answer the investigation question. Observation must be useful without becoming a new source of data exposure.
Alert on meaningful changes rather than every small fluctuation. A sudden increase in unauthorized tool requests could signal abuse or a prompt regression. A rising “no evidence found” rate might indicate that document ingestion failed. Higher token usage could reflect newly added context or a loop. Route alerts to owners who can act and document the likely failure path. Review whether the user journey remains successful after infrastructure recovery. A restored HTTP success percentage does not prove that agents are making correct decisions if the retrieval index remains stale.
A useful reliability drill recreates the difference between an apparently successful answer and an unsafe side effect. Consider a claims assistant that receives a tool timeout after requesting a reimbursement. The test should verify that the backend stores an idempotency key, that the assistant queries transaction state before considering a retry, and that the user is not told payment succeeded until the authoritative system confirms it. Run the same test when the tool returns malformed JSON, partial success, or a permission denial. These cases measure workflow integrity, not just language quality.
Evaluation ownership should be explicit. Product specialists judge usefulness, security engineers define prohibited effects, and platform engineers verify retry and recovery behavior. A failed safety assertion should block release even if the overall average user rating improves. Group defects by root cause: model misinterpretation, missing context, retrieval error, tool schema ambiguity or weak backend enforcement. Without that classification, a team may repeatedly alter a prompt to compensate for a permissions defect and leave the real failure mode intact.
Release changes with a recovery path
Version prompts, tool schemas, application code, retrieval rules and model configuration together in a deployable manifest. Before a release, compare the candidate against the approved baseline using the regression suite and a targeted set of new cases. Staged rollout or shadow evaluation can reduce exposure where architecture allows, provided user data handling remains compliant. Set rollback conditions before deployment. If the assistant begins making unauthorized tool requests, the release should be stopped even when latency and user sentiment are otherwise positive. Critical safety checks should not be averaged away.
After an incident, analyze the causal chain and add a regression test that would have caught the failure. Do not solve every issue with another vague instruction; sometimes the correct fix is stricter tool validation, a better retrieval filter or state persisted outside the model. A mature testing program makes future releases safer because each real defect becomes a specified behavior with a permanent owner. Developers preparing for CCDV-F should practice explaining that engineering boundary: generative flexibility is valuable, but reliable products require deterministic contracts, testable assumptions and accountable recovery.