WORKBENCH/ DISPATCH 007 / 7

Fourteen MCP servers to FastMCP 4 in two days

Why the framework upgrade was the easy part: audit first, safety net second, one pilot, then waves.

  • fastmcp
  • mcp
  • python
  • github-actions
  • cloudflare
STATUS · SHIPPED
Fourteen brass-and-walnut modules in two rows, each with a green light, beside a ticked checklist ending in "shipped". AI-generated with Bildsprache.

I expected the framework upgrade to be the hard part. It took the least time of anything.

Over Sunday and Monday, 27 and 28 September, the 14 MCP servers I run moved from FastMCP 3.x to 4.0.10. That’s 77 merged pull requests across 17 repositories, two servers retired, and one data exposure found and closed along the way. Most of it was done by Claude Code agents working in parallel; every merge went through me.

Here’s what made it fast, in the order it happened.

1. Audit before touching anything

Read-only agents went through the fleet in parallel, one per group of servers, plus one that read the FastMCP 4.0.x changelogs and the Cloudflare portal docs. Nobody was allowed to commit.

The useful surprise: the migration itself was almost free. Trial runs on 4.0.10 passed 401 of 401 tests in one server and 265 of 265 in another, with the camelCase compatibility switched off. What the audit did find was worse than a framework change. In four repos, CI ran pytest || echo "Tests skipped or failed" and then deployed anyway. One server’s tests had never run in CI at all, because pytest wasn’t installed there. Only one repo had branch protection.

So the plan changed shape. The framework moved to step four.

2. A safety net first

One pull request per server: a CI job that fails on red tests, a small protocol test (start the server, list its tools, make one real call through the in-memory client), then branch protection requiring that check. Releases stopped committing version bumps to main and push a git tag instead, which also ended a lockfile drift and a double deploy per merge.

Every one of those PRs was green before a single server changed framework.

3. Telemetry on the old version

A small middleware now writes one JSON line per tool call: server, tool, duration, outcome, and (on 4.x) the negotiated protocol. No arguments, ever. The hook is identical on 3.x and 4.x, so it shipped before the migration and a small script reads it back across four servers and the launchd logs on my Mac.

That line turned out to be the best test instrument I had. It showed that the Cloudflare MCP portal talks to my servers in protocol 2025-11-25, one version behind what my servers offer, and it showed failures that had been reported as successes.

4. One pilot, then waves

The pilot was ytdlp: rarely used, deployed like the rest, and it has long-running downloads, the same timeout problem my image server has. It got FastMCP 4.0.10 plus a job-and-poll pattern (return a job id after 20 seconds, fetch the result later), because background tasks don’t pass through the portal yet.

The pilot rewrote the recipe before anyone else used it. One finding: camelCase keyword arguments still work in v4 and only camelCase reads fail, so a grep won’t find them but the tests will. Then wave 1 (seven servers, in parallel), then wave 2 one at a time, including the three that run on my Mac under launchd and hold macOS permissions tied to one Python binary.

5. Verify the effect, not the signal

Every deploy got checked by what the container actually runs and by one real call through the portal. That habit found more than the migration did:

  • ytdlp had failed every download since May, because a volume stayed root-owned after a switch to a non-root user. The healthcheck was green the whole time.
  • A deploy trigger had returned 401 for three weeks while every merge looked shipped.
  • A private repo’s container image was public, and it held brand files. It’s private now, and that stack builds from source.
  • Six invoicing tools called API routes that don’t exist; their tests mocked those routes and passed.

The wiki now has a page for this pattern. I called it false success.

Why it was easy

Three reasons, and none of them is the framework. The audit ran before any change, so the plan fit the fleet instead of the changelog. The safety net went in first, so every later step had a test that could fail. And the tooling is my own: the text you’re reading was checked by my Klartext server, the images came from Bildsprache, and the draft went in through Writings.

The public servers, if you want to read the code:

The first real usage report is due on 12 October. Two servers are paused until it tells me whether anyone calls them.

If this poked at something you're building: talk to me.

BUILT.