Three talks that stuck with us from AGNTCon + MCPCon Europe
Two days of AGNTCon + MCPCon Europe at the RAI. Three talks that stuck with us: Booking.com's MCP gateway, Dexter Horthy on the software factory, and conformance testing for MCP.
On 17 and 18 September, AGNTCon + MCPCon Europe took place at the RAI in Amsterdam: two days about AI agents and the Model Context Protocol (MCP). Four of us went: Inico, Martijn, Jorick and Pascal. For a company that is all about testing and shipping, the topic is very relevant for us. When agents write code and call tools on their own, the question of how you know something works changes too. Three talks stuck with us.

One gateway for every client
Anushka Bhandari of Booking.com spoke about "From MCP Playground To Org-Wide Infrastructure: Lessons From Building Booking.com's Agent Foundry". Her starting point: most MCP projects die somewhere between proof of concept and production, and that gap is organizational. Who owns the servers? Who reviews contributions? And how do a designer, a product manager and an agent share the same infrastructure without ten different logins?
Booking.com's answer is the Agent Foundry: a two-tier MCP gateway with more than twenty org-wide servers, including Grafana, Honeycomb, Atlassian, Slack and GitLab. People and agents use the same OAuth flow. There's a skills registry where contributions are reviewed by AI, and profiles bundle MCPs and skills per workflow.
One of the slides puts it in a single picture. A developer picks their own client, whether that's Cursor, Windsurf, Claude Code or Codex. The gateway handles routing and authorization to the company systems, and the permissions that already apply there stay in place. Our colleague Martijn summed it up as "an MCP gateway that can connect clients and models to company systems".

Photo: Martijn
According to Martijn, two things make the approach so strong. The first is that the demand came from inside the company. People asked for it themselves: "I want to use that too!" As Martijn puts it, that makes it a solution that proved and sold itself, instead of a tool that was pushed onto people. Or, with a wink: built for developers, wanted by the whole company. The second is ownership. Whoever adds a new MCP also maintains it. You build it, you run it.
The software factory
In his keynote "State of the Software Factory" on day two, Dexter Horthy of HumanLayer showed how the loops in the software factory are shifting. That's what stuck with us most: what changes at each step, and what's possible now.
He starts in 2022, just before AI. Users send in complaints and feature requests, someone builds the thing, and then comes a pull request with checks and a human who tests and reviews the change. Building takes hours or days, and so does review. That's why teams plan up front. An hour of planning saves rework and makes the review shorter.

Photo: Pascal
Then comes the agentic software factory. An agent builds, and building drops from hours or days to minutes or hours. Review still takes hours or days and becomes the bottleneck. So agents are added to review code and run regression tests, and incidents and user feedback get routed into the factory as well. The question becomes how much you can put into the queue, and how fast you can review and test what comes out.
The step after that is the lights-off factory, where nobody reads the code anymore. The effort goes into tests, sandboxes, automated review, monitoring, rollout and user feedback. According to Horthy, that doesn't work yet, because verifying quality is much harder than checking whether the tests pass.
His answer: turn the lights back on. Code review comes back, and you plan up front, with help from AI, in four phases: product, system architecture, program design and vertical slices. "30 minutes of planning saves hours of review." Better to go two to three times faster safely than ten to a hundred times faster.
He expands on this in his essay "Why Software Factories Fail". One line from it: "Tests give you feedback in seconds, but the cost function of bad architecture is measured in weeks, months, maybe even years."
Knowing what your client really supports
Paul Carleton and Felix Weinberger of Anthropic spoke about "MCP Conformance Testing V1.0, Testing the 2026-07-28 Spec in SDK's and Online". Conformance tests check whether SDKs implement the MCP specification in a way that lets them work together. The 2026-07-28 spec revision is the first one where those tests are a required part of the process for proposing changes to the spec. They also introduced hosted conformance testing, which clients and servers can use to test their deployed implementations.
The slide asks: "What if one page showed what your client really supports, per spec version?" You start a run, paste one config into your client and read one table, with 48 client scenarios per spec version. The run on that slide used Claude Code, on 16 September 2026. Want to try it yourself? Go to mcp-c9e-v1.val.run.

Photo: Pascal
It takes us back to the days when we checked whether everything worked in every browser. With MCP we're getting that again, because not every client supports the same parts of the spec. Checking conformance up front keeps you from finding that out in production.
What we take away
Keep ownership in mind, the way Booking.com does: developers pick their own client, and whoever adds an MCP also maintains it. Dexter gave a clear picture of what already works, and of which loops you can automate. And Paul and Felix showed that an MCP you build is easy to check against the spec.
Want to get started yourself? Take a look at our training.
Ready to ship faster?
Book a free 30-minute intro call.