Skip to main content
Viithiisys
Back to blog
AI Engineering6 min readTue, Oct 06, 2026

AI Agent Development Services and the MCP Risk

MCP trust gaps let one compromised agent steer others. What AI agent development services should build in: egress limits, scoped credentials, zero trust.

Gaurav Saini

Founder, Viithiisys

AI Agent Development Services and the MCP Risk

What is protocol pivoting in MCP agent networks?

Protocol pivoting is an attack where malicious instructions enter one agent through MCP and are forwarded to another agent over a different protocol, such as A2A. The receiving agent trusts the sender, so it acts on the instruction.

Anyone evaluating AI agent development services should understand this pattern, because it targets the agent rather than the model.

According to Ars Technica, independent researcher Syed Anas Mohiuddin tested agents from Google, JP Morgan Chase, Weviate, Rapid7, the French government's interministerial digital directorate and the US federal government. Google and four other organisations acknowledged vulnerabilities of this kind within five months.

Markus Vervier of X41 D-Sec argues the better name is indirect prompt injection, with pivoting as a subclass. The label matters less than the consequence: text planted in content ends up as a trusted task inside your network.

Why does agent-to-agent trust fail?

Agent-to-agent trust fails because MCP servers store credentials for each agent, and agents are built to trust every other internal agent. An instruction an LLM would refuse can succeed when it arrives as delegated work.

Where did the guardrails go?

Many special-purpose agents, such as a translation or data-analysis agent, have thin guardrails or none. Their job is to process input and pass results along, so they forward whatever they receive.

Rapid7's Douglas McKee described the failure to Ars: every component did what it was designed to do. "Each protocol was built assuming it lived on its own, so each one checks its own front door while nobody watches the hallway in between."

What did the Google and Rapid7 cases look like?

Both were flaws in how a component handled requests it should not have trusted. Google's case was a redirect-following bug; Rapid7's was a low-scored flaw that still mattered in a chain.

Google's MCP toolbox for databases (googleapis/mcp-toolbox) initialised its HTTP client without a CheckRedirect policy and did not validate target IP addresses. A crafted path parameter could make it follow a redirect to an internal endpoint. The issue was rated 8 out of 10, and Google fixed it with IP allow-lists and block lists.

Rapid7's CVE-2026-97228 scored only 2.7 out of 10 and was fixed last month. Our reading: a low score on one component understates the risk once that component sits in a chain of agents.

Is MCP itself the problem?

MCP is not the flaw in isolation. The weakness is the assumption that everything inside the network is friendly. The Model Context Protocol standardises how apps and agents reach tools, and it does that job.

Agent networks rebuilt the flat, implicitly trusted internal network that zero trust was designed to retire.

Zero trust, as described in NIST SP 800-207, assumes any node may already be compromised and verifies each request. Ars notes that organisations rushing to build agentic architectures have largely dropped that principle.

The trade-off is real. Verifying every agent-to-agent call adds latency, identity plumbing and logging work. Teams skip it for the same reason they once skipped internal TLS: it slows the demo.

What controls stop a pivot from spreading?

Four controls cover most of the exposure: narrow per-agent credentials, egress allow-lists on MCP servers, validation of inter-agent messages, and logging of delegation chains.

None of them needs new research. Each is ordinary engineering that a team can scope, price and schedule, and each removes a different link in the pivot chain.

The list below sets out what each control stops and what it costs, so you can weigh the trade-off before committing.

  • Per-agent, task-scoped credentials. Stops one compromised agent reaching every connected system. Cost: more credential management and harder local debugging.
  • Egress allow-list and redirect policy. Stops SSRF and calls to internal endpoints. Cost: each new integration needs an explicit rule.
  • Inter-agent messages treated as untrusted. Stops injected instructions passed down the chain. Cost: an extra validation step and occasional false rejections.
  • Delegation logging with origin tracking. Stops slow, silent exfiltration you cannot trace. Cost: log volume and storage, and someone must read it.

How should MCP servers handle outbound requests?

Copy the pattern Google adopted. Apply an allow-list of IP ranges, block internal and metadata addresses, set an explicit CheckRedirect policy, and reject an unsafe base URL at startup rather than on the first request.

Researcher Syed called this "more work than most MCP servers have done." It is a few days of engineering, not a research project.

How do you limit what a delegated task can do?

Give every agent its own credential with the minimum scope for its job. A translation agent should never hold database write access.

Then carry the origin of each task with it. If an agent receives work that traces back to untrusted content, such as an inbound email or a scraped page, it should run with reduced permissions.

What should you ask AI agent development services vendors?

Ask for the threat model before the demo. A credible vendor can describe which agents trust which, what each credential can touch, and what happens when one agent is fed hostile text.

Specific questions worth putting in writing:

  • Which protocols do the agents use, and where does one hand off to another?
  • What outbound network access does each MCP server have?
  • Who reviews the architecture before production, and who owns incident response?
  • What is logged between agents, and for how long?

This applies equally to generative AI development services, workflow automation and chatbots, since any of them can end up as a node in an agent chain. It matters most in regulated work. Healthcare software development services handle patient records, where the exfiltration risk Ars describes is a compliance event, not just a bug.

How do we approach agent security at Viithiisys?

Viithiisys has been shipping software since 2007 from Mohali, with a Canadian office in Markham, Ontario.

Across that time, the lesson that repeats is that internal systems get compromised through the trusted path, not the front door. A service account that can reach everything, or an internal tool nobody reviewed, is where incidents start. Agent networks make that old pattern faster and harder to see.

In our AI agent development and agentic AI development work, we draw the trust boundaries on a diagram before writing integration code. Each agent gets its own credential, and each MCP server gets an explicit egress list.

We will not claim this is complete. Prompt injection has no full fix today, so the design goal is to limit what a successful injection can reach.

Does this change how you scope AI agent development services?

Yes: scope the first version so the attack surface stays small. A single agent, one protocol and read-only access to one data source removes the pivot path entirely, and you can add delegation once the controls exist.

That is how a Moonship MVP works. It is built around a 30-day delivery window against a fixed scope, and the window is realistic because the scope is small: one agent, one protocol, one data source.

A fixed scope also forces an early decision about which agent talks to which system. It is a sound pattern for any team buying MVP development services or custom software development services for an agent product.

If your team has no one senior enough to own these decisions, a fractional CTO can review the architecture on a fractional basis. If agents are already live, book a broken workflow assessment and we will map where trust is implicit. For a scoping conversation first, use the contact form.

FAQ

What is protocol pivoting in AI agents?
Protocol pivoting is a multi-step attack where an adversary gains access through one agent protocol, such as MCP, then exploits trust assumptions to reach capabilities exposed through another, such as A2A. Researcher Syed Anas Mohiuddin coined the term. Some security researchers class it as a subtype of indirect prompt injection.
Is MCP unsafe to use in production?
MCP is usable in production, but its servers often hold credentials and few ship with SSRF guards or redirect policies. The risk sits in how teams deploy it: broad credentials, implicit trust between agents and no egress limits. Scope each credential narrowly and validate every outbound request.
What should I ask an AI agent development company about security?
Ask how credentials are scoped per agent, whether inter-agent messages are validated as untrusted input, what outbound network access each MCP server has, and how redirects and internal IP ranges are blocked. Also ask who reviews the agent architecture before launch, and what gets logged between agents.
Do small agent projects need zero trust?
Yes, in proportion. A single-agent MVP with one narrow credential and no internal network access needs little machinery. The moment a second agent can delegate work to the first, assume either may be compromised and give each only the access its own task requires.