End-to-end: schema, OAuth-style auth, scoped tool surface, structured errors, observability hooks, deployment to a single VPS, and registration in Claude Code and Cowork. Skip the official Python SDK template for production: TypeScript with Hono cuts deploy time in half and dodges the SDK's known RCE class.
Most internal MCP servers we are asked to look at began as a weekend build from the official SDK template. It worked in the demo, so it was pointed at the CRM or the warehouse database, defaults included. The template was written to teach the protocol, not to sit between a language model and your production data.
This is the build guide that precedes our MCP server hardening checklist. The checklist says what a production server must have; this is how we build one that has it from the first commit: seven steps, TypeScript with Hono, a single VPS, registered in Claude Code and Cowork.
One framing makes the rest of the decisions easy. The client of an MCP server is a language model, and a language model can be talked into calling anything you expose by anyone who can get text in front of it, including the author of a support ticket.
Why TypeScript and Hono rather than the Python SDK template?
Two reasons. The first is the SDK's known RCE class. The disclosed vulnerabilities were a class rather than a single bug, and the class lives in the layer the template does for you: taking bytes from the transport and dispatching them to a handler with more trust than a network boundary deserves. With Hono you own that layer: a small router, explicit parsing against your schema, explicit dispatch to a named handler, nothing else, and every line of it is code you wrote and can audit.
The second is deploy time. A TypeScript build is one artefact running as one process on a runtime you already have, and in our experience that cuts deploy time in half against the Python template: the saving is in steps that no longer exist. The honest trade-off is that the Python template is the quickest route to a demo, and a Python-only team pays something to learn Hono. Worth paying, because a security boundary should be built from parts you understand.
What are the seven steps?
- Write the schema first. Define every tool's inputs and outputs as a typed schema before any handler exists. The schema is the contract the model reads, the validator the server enforces and the document your reviewers sign off; a field that is not in it never reaches a handler.
- Add OAuth-style auth. Your identity provider issues short-lived bearer tokens, the client carries them in the Authorization header, and the server validates each request before it parses the body. Nothing is accepted from a query string, and the token's subject travels with the call so each tool acts as a person rather than as the server.
- Scope the tool surface. Expose the minimum set of tools the job needs, named as verbs with narrow arguments, with reads and writes as separate tools carrying separate permissions. Default to deny, and add a tool only when someone can name the workflow that needs it.
- Return structured errors. Every failure comes back as a typed error the model can act on: not authorised, not found, invalid argument, rate limited, upstream unavailable. A stack trace in a tool result is a leak and an invitation to improvise; a typed error lets the model retry what is retryable and report the rest.
- Wire the observability hooks. Every call is logged with its tool, arguments, subject, session, latency, result size and outcome; the per-session rate limit lives in the same hook. Emit to whatever you already run; this is the audit log the checklist asks for, and it is cheaper on day one than after the incident.
- Deploy to a single VPS. One machine, one process under the operating system's service manager, a reverse proxy terminating TLS, secrets injected from the environment, dependencies pinned, and egress restricted to the internal tool and the identity provider. A machine you understand beats an orchestrator you do not for a server handling one team's calls.
- Register it in Claude Code and Cowork. Add the server to Claude Code as a remote HTTP server with the auth flow, and to Cowork as an organisation-level connector, so the approved-tools page in your AI policy points at something that exists. Smoke-test every tool from both clients before anyone else gets the URL.
How should auth work for an internal MCP server?
OAuth-style means the server is a resource server and nothing more: it issues no tokens, it only checks them. The subject and scopes in the token drive authorisation per tool, so `find_invoice` runs as the person who asked, with the permissions they already have.
The consequence we push hardest is that the server should hold no master credential to the internal tool: act on behalf of the user where the tool supports it, and otherwise hold a read-only credential for the read tools and a separate one for the writes, each scoped to the call. Where the internal tool has no per-user permissions at all, the server becomes the authorisation layer and checks the subject against an allowlist per tool before the handler runs. More work than a shared key in an environment variable; also the difference between an incident touching one person's access and one touching everything.
What does a good tool surface look like?
Small, verb-named and narrow. Tools map to workflows rather than to the database: `find_invoice` with a customer and a date range, not `run_query` with a string. The description is written for the model, not the developer: what the tool does, when to use it and what it will refuse to do.
Then apply the checklist to what comes back. The model cannot tell a tool result from an instruction embedded in it, so anything a user typed is sanitised before it leaves the server; results are capped and paginated; every write tool has a dry-run mode or a confirmation step. The test before a tool ships: could a prompt injection in a ticket body do damage through this surface? If yes, it is too wide.
An MCP server is an API whose client can be talked into calling anything you expose. Build the surface for that client, not for the demo.
How do you register it in Claude Code and Cowork?
Claude Code takes the server as a remote HTTP server in its MCP configuration, at project level so the team shares one definition; the OAuth flow prompts a sign-in the first time a tool is called, and short-lived tokens keep the local token store a small problem. Cowork takes the same server as an organisation-level connector, so the operations person who never opens a terminal gets the same tools under the same auth. One server, two clients, one log is the point.
Registration is also where the approved-tools page of your AI usage policy stops being a PDF. The server is the approved path to the internal tool, and adding a tool to it is a change with an owner, a schema and a log line. Write down who can add the next one.
What to do next
If you have a template server in production, walk the hardening checklist this week and plan the rebuild on these seven steps; schema and auth are the two you cannot retrofit cleanly, so start there. If you are about to build your first, write the schema before the first handler.
This is the kind of work we do on spec: a 20 or 45-minute call about which internal tool Claude needs to reach and who should reach it, then, where the scope needs defining, a four-day Spec from €5,000 that turns the answer into a written brief you own outright. We have no product to sell you. The server is yours: on your VPS, in your identity provider, with your log.