Skip to content

Blog Article

MCP Testing: A Practical Server Test Ladder

Test an MCP server from tool discovery through invalid input, result shape, denied writes, authentication boundaries, and recovery with a bounded local fixture.

By AppHandoff Team · Published · 11 min read

Illustration of an MCP connection linking an agent with shared project context
MCPTesting

MCP testing works best as a ladder: prove the client can discover the right tools, reject bad arguments, return the intended result shape, deny unauthorized writes, respect identity boundaries, and recover after state or credential changes. A green tool call is only one rung. It does not prove that the application returned the right project record or made the requested state change. This guide gives a bounded fixture you can run locally, then shows what to verify against a real server.

The protocol details here follow the MCP tools specification, revision 2026-07-28, and the official MCP Inspector guide. The example names and records below are invented for illustration. The local handler fixture was run for this article; it is not an MCP server and does not establish protocol conformance.

Define the evidence before running a tool

Write down one question per layer. Can the intended client reach the endpoint? Does its discovered catalogue contain the expected tool and input schema? Does malformed input fail without changing state? Does a valid call return the right record and the fields a consumer actually needs? Does a caller without write permission get a refusal while the record remains unchanged? Does a second identity fail to read the first identity's project? A single “tool succeeded” assertion cannot answer those questions.

Choose disposable records and credentials with no customer data. Give the fixture an initial state, an exact expected output, and a cleanup method. If the server can write, test its refusal path before testing an authorized write. Record the tested protocol revision, client version, transport, identity, scope, project, and server build. That context makes a failure reproducible without pasting secrets into a report. For a remote server, a read-only identity is a sensible first connection.

Rung 1: connection and discovery

Start the server using its documented transport and connect with the client you plan to support. In the current MCP specification, servers with a tools capability answer tools/list with tools available to the requesting client. Tool definitions include a name and an inputSchema; some also provide an outputSchema. The catalogue can depend on authorization presented with the request. Assert the expected name, meaningful description, and argument schema rather than merely counting tools. If pagination is present, follow its cursor before declaring a tool absent.

Use MCP Inspector to connect and inspect a real server's catalogue and responses. A 200 from a health URL proves reachability of that URL, not MCP negotiation or tool discovery. A client configuration that parses also leaves the connection untested. The MCP config checker checks selected core fields of strict JSON for Claude Code, Cursor, or VS Code in the browser; use placeholders for credentials. It does not launch a process, contact an endpoint, perform OAuth, or certify full client compatibility.

Rung 2: validation and input boundaries

Call one safe read with a valid record ID and then with a wrong type, a missing required field, and a well-formed but unknown ID. The first should return the known record. The others should produce distinct, understandable failures where the application needs that distinction. Test actual runtime validation: publishing an input schema is useful to clients, but the server must still reject invalid arguments before a handler performs a write. Verify the data store stayed unchanged after each rejected request.

Keep the test set small and deliberate. An ID of the wrong type probes validation; an unknown but well-formed ID probes lookup behavior; an ID from another project probes authorization. Treating all three as one generic error hides which boundary failed. If an argument is optional, test omission and a supplied value. If the tool accepts pagination or a byte limit, assert a bounded response and a usable continuation. Those cases matter most when another agent will consume the output without a human watching each call.

Rung 3: result shape and application truth

Inspect the complete tool result, including isError, text content, and structuredContent when the server provides it. The current tools specification permits structured results and says a declared output schema must match them. A client should validate a structured result against that schema. Your application test should go further: assert the exact record ID, project, status, revision, and any field the next action depends on. A syntactically valid response for the wrong record is a failure.

For write tools, compare state before and after the call using an independent read. A successful response that says “closed” while the stored record remains open is not a completed workflow. Conversely, a tool error that already changed state is a serious failure. Account for eventual consistency if the product documents it: use a bounded retry with a deadline and report the observed delay, rather than immediately treating a stale read as success or sleeping indefinitely. The MCP server architecture guide explains where the protocol boundary ends and the application record begins.

Rung 4: a disposable handler fixture

Here is the small shape behind our local check. It intentionally tests a handler, not HTTP, stdio, JSON-RPC, OAuth, or a complete MCP tool definition. The invented read_work_item and close_work_item names are not AppHandoff tools. The result objects are illustrative handler outputs, not complete modern MCP wire messages.

import assert from 'node:assert/strict';
const records = new Map([['T-1', { id: 'T-1', status: 'open' }]]);

function call(name, args, context) {
  if (name === 'read_work_item') {
    if (typeof args?.id !== 'string' || !/^T-\d+$/.test(args.id))
      return { isError: true, code: 'INVALID_INPUT' };
    const record = records.get(args.id);
    return record
      ? { isError: false, structuredContent: { ticket: { ...record } } }
      : { isError: true, code: 'NOT_FOUND' };
  }
  if (name === 'close_work_item') {
    if (!context?.canWrite) return { isError: true, code: 'WRITE_DENIED' };
    if (typeof args?.id !== 'string' || !records.has(args.id))
      return { isError: true, code: 'INVALID_INPUT' };
    const next = { ...records.get(args.id), status: 'closed' };
    records.set(args.id, next);
    return { isError: false, structuredContent: { ticket: next } };
  }
  return { isError: true, code: 'UNKNOWN_TOOL' };
}

assert.deepEqual(call('read_work_item', { id: 7 }, { canWrite: false }),
  { isError: true, code: 'INVALID_INPUT' });
assert.deepEqual(call('close_work_item', { id: 'T-1' }, { canWrite: false }),
  { isError: true, code: 'WRITE_DENIED' });
assert.equal(records.get('T-1').status, 'open');
assert.deepEqual(
  call('read_work_item', { id: 'T-1' }, { canWrite: false }).structuredContent,
  { ticket: { id: 'T-1', status: 'open' } }
);

Save that code as an .mjs file and run it with Node. The assertions should exit without an error. Our disposable copy passed invalid input, denied write, unchanged-state, and bounded read-shape assertions. If the denied call accidentally mutates the map, the unchanged-state assertion fails. If the read returns a plausible but wrong ticket, the final deep equality fails. The example deliberately does not simulate authentication; a Boolean permission flag cannot replace a real session, token, or project membership check.

Rung 5: rule denials and identity boundaries

Run a real denied-write check with an identity that lacks the action's permission and a disposable target record. Assert the refusal code or documented error shape, the absence of a state change, and the remedy offered to the caller. Then use an authorized test identity on a separate disposable record, if that action is approved for the test environment. Never infer authorization from tool visibility. Discovery tells a client what may be called; the server's action rules decide what this caller may do to this record.

Next test tenant or project isolation with two controlled identities. Give each identity an unambiguous record. A caller from project A must not read or write project B's record by guessing its ID. Verify both the returned result and the underlying state. For protected HTTP servers, also test no token, expired token, wrong audience, and insufficient scope as separate cases. The current MCP authorization specification defines the HTTP authorization framework; it does not prescribe your product's project membership or approval rules.

Rung 6: recovery after change

After a clean read, change the disposable record or its revision and repeat the call. The server should return the new state or a documented conflict, not silently report the old result as current. Revoke or expire a test credential and confirm that the next protected request is refused. Refresh or replace it through the documented flow, then repeat a permitted read. Keep retries bounded so a broken refresh or a persistent denial cannot become an endless loop.

Also test a client that has cached an older tool catalogue: reconnect or refresh discovery after a server release, then confirm that the client sees the current names and schemas. If a tool was removed, an old invocation should fail plainly. If a rule revision changed, fetch current guidance before deciding whether to retry. Recovery evidence is the new successful read or explicit refusal, not merely a reconnect button click. Record the failing layer and one corrective action so the next engineer can distinguish transport, auth, tool input, and business rule problems.

Applying the ladder to AppHandoff

AppHandoff's documented endpoint is https://api.apphandoff.com/mcp. Its current server exposes eight tools: bootstrap, get, find, ticket, plan, message, project, and decide_lifecycle_proposal. The last serves a signed-in human approval card and is not listed to clients without that support. Its connection guide helps choose a client setup. For a safe first check, discover tools, call bootstrap for a project your account can reach, then use get or find for a named read. Compare the project, ref, and returned content with what you expected.

The product's write rules add scope, revision, and sometimes human approval requirements. A tool call that needs a human lifecycle decision yields an approval path; an agent must wait for the signed-in person. Do not use an approval card action as an automated test of agent write permission. For a denied write in a controlled test, inspect the specific structured refusal and verify no state change. The current MCP versus API guide explains why a discoverable tool does not replace the application's authorization contract.

Keep the test report honest

A useful report lists each rung with the exact environment, request class, expected outcome, observed outcome, and the record or response evidence. Separate “fixture passed” from “server passed.” The local fixture above proves four handler assertions in isolation; it says nothing about transport, current protocol metadata, OAuth, real AppHandoff behavior, or every client. A real Inspector session can prove discovery and a permitted read for the chosen identity and environment, but it still cannot prove every application workflow. Expand the suite where failures would be costly, especially cross-project access and writes.

For the next iteration, turn repeated production defects into narrow regression cases: the original input, the rule that should reject or accept it, and a post-call state assertion. Keep credentials and customer data out of fixtures. Re-run the ladder when tool schemas, auth rules, or result shapes change. That is the practical value of MCP testing: evidence for each boundary, with clear limits on what the evidence establishes.

Frequently asked questions

How do I test my own MCP server?

Start with the server and client version you intend to support. Confirm connection and tool discovery, then call one safe read with valid and invalid arguments. Check the complete result shape, denied writes, authentication boundaries, and recovery after a changed record or expired credential. Keep protocol tests separate from application rule tests.

Does a successful MCP tool call prove my application works?

No. A successful tool call proves only that this request reached a handler and received a response. The response can still contain the wrong record, omit a required field, leak another project's data, or leave an intended write unapplied. Assert the returned data and the application state independently.

Can a local handler fixture establish MCP conformance?

No. A handler fixture can exercise validation, permission decisions, and result mapping. It does not test a transport, the current request metadata, discovery, authorization, client behavior, or the full protocol. Use a real client or MCP Inspector against a disposable server for those checks.