The cassette is itself an MCP server, so any client in any language connects to it exactly as it connects to the live one: no library to import, no product code to change, no transport to wrap.
That sentence is one line to write and considerably more than one line to implement. This is what it costs: an incremental parser for a streaming format whose line endings are not obvious, two protocol lifecycles in one file, and a header that cannot be written when the file is opened.
A test suite for an MCP client needs a server. There are two usual ways to get one, and both charge you.
Run the real server. Correct by construction, and it costs what real things cost: credentials in CI, rate limits, flaky network failures you did not schedule, state that survives runs, and a counterparty free to change its answers without telling you. Tolerable once; not on every push.
Write a fake and run it in the same process. Fast, hermetic, no credentials, and two properties easy to miss while writing it. It is a library for one language, so a team with clients in two languages writes and maintains two fakes. And it is assembled from your reading of the specification rather than from what a server put on the wire, with nothing keeping the two in step. It passes because it agrees with you.
Before deciding which parts of this project were worth keeping, we measured the population it serves: 158 public MCP servers, read by hand where the automated pass was ambiguous. Of those, 83 test through the protocol under the strict reading and 89 under the looser one, which is 52.5% and 56.3%; the honest bracket after accounting for both measurement errors found is 45 to 60%. Among those 83, 16 (19.2%) reach for an official SDK's in-memory transport to avoid live APIs. The method, including two measurement defects that moved the number and were reported rather than quietly fixed, is in docs/research/01-reality-check.md.
Read that the way it falls: protocol-layer testing is already a mainstream habit, not a frontier, and most of that majority got there without this tool. The open question is what sits at the other end, and whether it is the same artifact for every language you write clients in.
The recorder is a transparent stdio proxy. Point your client at it instead of at the server, give it the server command, and it forwards bytes verbatim in both directions while writing every JSON-RPC frame it recognises to an append-only JSONL file (src/record.ts). Neither side can tell it is there, and what lands in the file is what the server said rather than what a wrapper decided it meant. Output that is not JSON-RPC at all, such as a chatty server's stdout logs, is preserved as raw entries, so the transcript stays faithful even to a misbehaving server.
Replay reads that file and serves it as an MCP server on stdio (src/replay.ts). Requests are fingerprinted and matched against recorded pairs: lifecycle calls by method, tools/call by tool name plus stable-stringified arguments, everything else by method plus params with the volatile parts (_meta, pagination cursors) removed. A fingerprint miss falls back to the next unconsumed response for the same method in recorded order, so a reordered test run still lands. A true miss reports the closest recorded fingerprint and exactly which component diverged.
The claim in the pitch is testable, so here it is tested: the same health check against a live reference server, then against the recording of it.
server: mcp-servers/everything@2.0.0 protocol: 2025-06-18
surface: 13 tools, 7 resources, 4 prompts
[OK] no findings
result: PASS (0 error(s), 0 warning(s))
Those four lines are the output twice: once with the real server running, once with nothing but a file. Now the part a fake in your own test suite cannot do, using the official Python SDK, which has never heard of this project:
SERVER = StdioServerParameters(
command="npx",
args=["-y", "mcp-cassette@0.4.0", "replay", "everything.cassette.jsonl"],
)
async with stdio_client(SERVER) as (read, write):
async with ClientSession(read, write) as session:
init = await session.initialize()
tools = await session.list_tools()
server: mcp-servers/everything 2.0.0
protocol: 2025-06-18
tools: 13
first three: ['echo', 'get-annotated-message', 'get-env']
One cassette, two clients in two languages, identical answers. The Python side imports nothing from here: it launches a subprocess and speaks stdio, exactly as it would to the real server. The rest of this page is what that costs to make true.
Over HTTP, an MCP response may arrive as a stream, and recording it faithfully means parsing it the way the specification says rather than the way it looks. The parser is src/sse.ts, and it implements WHATWG HTML's "interpreting an event stream" and nothing else. The rules that a hand-rolled split would get wrong:
if (c === "\r" && i === this.buffer.length - 1) break;. At end of stream, a held CR can only have been a line ending after all, and is processed then.data: x carries a leading space in its value.data line is not an event. An empty data: line is a different thing and does dispatch.id containing NUL is ignored, and the previous last event id stands.The chunk-boundary rule is the one with teeth: it fails only under timing you do not control and cannot be reproduced by reading the code. So it is tested by exhaustion rather than by example. One stream carrying all three line endings, a comment, a multi-line data field and a trailing bare CR is fed at every read size from 1 byte to the whole string, and every size must produce the same three events (tests/sse-capture.test.ts, "splits the same events no matter where the reads fall").
The protocol has two eras alive at once, and a recording has to know which one it captured.
The classic lifecycle (revisions up to 2025-11-25) opens with initialize, then notifications/initialized, then requests, and the server states its protocol version in that result. The 2026-07-28 revision is stateless: there is no handshake, every request carries _meta, and server/discover answers what initialize used to. Sessions are gone in that world, and so is the standalone GET stream.
A cassette records the era it saw in the header line, and replay serves in that era rather than guessing from frames (src/cassette.ts, src/http-replay.ts). The migration rule is one function, and worth stating out loud rather than leaving to a reader of the type declaration:
export function cassetteEra(header: CassetteHeader): Era {
return header.era ?? "legacy";
}
A cassette with no era field is legacy by definition. That is not a fallback chosen for convenience; it is what keeps every file recorded before the field existed working, with no migration step. The stdio recording made for this page is one of them: its header carries no era, and both clients above were served, correctly, as legacy.
Which leaves an ordering problem. The header is line 1 of an append-only file, so it is written before anything else. But the era is decided by the first successful exchange, which has not happened when the file is opened, and cannot be inferred from the first request: a dual-era client may try server/discover, get an error, and fall back to initialize. Guessing from that attempt records a lie.
So the HTTP recorder opens the writer in deferred-header mode: entries are buffered in memory, and setEra() writes the header and flushes everything behind it once an exchange has actually succeeded (src/cassette.ts, src/proxy.ts). Success is the whole test, not whether the method was a lifecycle one. A session where nothing succeeded flushes at close with the era omitted, which the reader treats as legacy by the rule above, and changing the era after the header has reached disk throws rather than pretending. The case easiest to get wrong is tested: a server/discover answered with a JSON-RPC error over a stream leaves the header with no era, and still records the failure.
Three limits, stated as they are in the machine-readable index:
verify for that.mcp-cassette replay: 1 server-initiated frame(s) in the cassette are not replayed in v1
The first two are consequences of the design rather than gaps in it. A recording is a statement about one past session, and treating it as one about the present is the mistake this tool makes easy, which is why verify re-fires recorded requests at a live server and diffs the answers.
The third is a gap, and the words in v1 are the admission. A server that samples, elicits, or notifies speaks a direction of the protocol this replay does not reproduce: the frames are in the file, and nothing plays them back. Test a client that depends on that traffic against a live server.
Node 20 or later, and network access for the first one only.
npx -y mcp-cassette@0.4.0 check --stdio "npx -y mcp-cassette@0.4.0 record -o everything.cassette.jsonl -- npx -y @modelcontextprotocol/server-everything@2026.7.4 stdio"
npx -y mcp-cassette@0.4.0 check --stdio "npx -y mcp-cassette@0.4.0 replay everything.cassette.jsonl"
npx -y mcp-cassette@0.4.0 lint everything.cassette.jsonl
The first drives the recording proxy with the health check, so the check is the client and you need none of your own. The second runs the identical check with no server and no network. The third reads the finished file against its own header:
everything.cassette.jsonl: header and frames agree
Every version here is pinned, and not only because the specification is moving. An unpinned npx will reuse a copy already in the local cache instead of resolving the registry: on the machine these commands were run, bare npx mcp-cassette ran 0.1.1 while latest was 0.4.0, silently, with different output. Pin, or install once and call the binary.
From there, point whatever client your tests already use at mcp-cassette replay <cassette>, in whatever language, and change nothing else.