How to Test an API When Your OpenAPI Spec and Your Code Have Drifted

An OpenAPI spec is supposed to be the source of truth. In practice, on a codebase where an AI coding agent is actively shipping backend changes, the spec is often a snapshot of what the API used to look like, not what it returns right now. Testing against the spec instead of the running API means testing against a description that's already partly wrong.
Why Specs Drift Faster Than Anyone Notices
A spec drifts for reasons that have nothing to do with anyone being careless. An AI coding agent adds a field to a response and doesn't think to update the YAML file, because updating the spec wasn't part of the task it was given. A refactor changes a status code from 200 to 201 for a newly-created resource, technically more correct, but the spec still says 200. A serializer applies a naming convention, camelCase instead of snake_case, that the spec was never written to reflect in the first place.
None of these show up as an error anywhere. The code runs, the endpoint responds, and the only place the truth and the description disagree is in a document nobody is actively watching.
What Happens When Tests Are Generated From the Spec
Tests generated directly from an OpenAPI document inherit whatever the document says, including the parts that are wrong. If the spec says a field is required and the running API actually made it optional in a recent change, the generated test enforces a rule the API no longer follows. If the spec is missing a field the API now returns, the generated test simply never checks it, because it doesn't know the field exists.
This produces a specific, frustrating failure pattern: tests that are internally consistent with the spec and completely disconnected from what the API actually does. A green suite doesn't mean the API is correct. It means the API matches a description that itself might not be correct.
Testing the Running API Instead of the Document
The fix isn't a better spec-parsing tool. It's testing something other than the spec: the actual, running API, observed directly.
TestSprite's Backend Testing 2.0 works this way. Rather than generating assertions from an OpenAPI document, the agent calls each real endpoint and records what actually comes back, actual field names, actual status codes, actual response shapes, and builds test coverage from that observation.
Other verification tools read your code and guess. TestSprite opens your app and uses it.
A spec, if one exists, can still be useful context for understanding intended structure. But the assertions themselves are grounded in what the API does, not what a document from three sprints ago says it should do. When the two disagree, the observed behavior is what gets tested, and the disagreement itself becomes visible as a specific, checkable finding rather than a silent gap.
A Scenario: A Procurement Platform Where the Spec Lied About a Field Type
A team building a B2B procurement platform maintains an OpenAPI spec for its purchase order approval API, originally written when the endpoint returned a simple approver ID as a string. Since then, an AI coding agent extended the endpoint to support multi-level approval chains, and the approver field now returns a structured object containing the approver's ID, role, and approval sequence number, not a plain string.
Nobody updated the spec. It still documents approver as a string.
A developer runs TestSprite against the live API rather than regenerating tests from the outdated document. The exploration agent submits a purchase order, checks the approval response, and observes the actual structure: an object, not a string. Because the test is built from that observation, it correctly validates the object's shape, the approver ID, role, and sequence number are all present and correctly typed, rather than failing on a type mismatch the way a spec-derived test would, or silently passing a shallow check the way a loosely-typed spec-derived test might.
The developer also gets a specific note in the results: the observed response shape for approver doesn't match the type declared in the spec file, with the actual structure documented for reference. That's useful independent of the test itself, since it flags exactly where the spec needs to be updated, without anyone having to manually diff the document against the API by hand.
Deciding When the Spec Needs to Be the Source of Truth Anyway
There are cases where testing strictly against a spec is the right call, particularly for a public API with external consumers who depend on the documented contract regardless of what the implementation currently does. In that situation, a deviation from the spec is the bug, even if the running code is internally consistent.
The distinction that matters is who the API serves. For an internal API consumed only by your own frontend or by services you control, testing the observed behavior directly is almost always more useful, since the goal is confirming the whole system works together, not confirming compliance with a document. For a public API with external integrators, spec compliance itself becomes a requirement worth testing directly, and a deviation should be treated as a contract break regardless of whether the new behavior seems more correct.
Keeping the Spec and the Reality From Drifting Further Apart
Testing the running API catches drift after it's happened. Preventing drift from compounding requires treating the gap TestSprite surfaces, observed behavior against documented behavior, as something worth acting on, not just noting.
Running this check as part of the same trigger used after an AI coding session, "Help me test this project with TestSprite," through the MCP Server inside Claude Code or Cursor, means the drift gets caught in the same session where the change was made, while the context for whether the new behavior or the old spec is correct is still fresh.
Conclusion
An OpenAPI spec is a description of an API, not the API itself, and descriptions go stale in ways nobody notices until a test built from the wrong assumption either fails for the wrong reason or passes when it shouldn't.
TestSprite tests what the API actually returns, flags where that diverges from the documented spec, and leaves the judgment of which one is correct to the team, rather than assuming the document is always right.
Test your real API with TestSprite and stop finding out your spec was wrong from a production incident.