AI Systems Architecture

Connecting AI to Enterprise Systems with MCP: Architecture, Control and Risk

How to connect AI applications to enterprise operations without weakening the controls that keep live systems safe.

MCP gives AI applications a standard way to discover data and operations. In an enterprise platform, the harder work begins after the connection has been made.

Published: 2026 14-16 min read AI Systems Architecture
  • MCP
  • AI Systems Architecture
  • Enterprise Integration
  • Operational Control

MCP gives AI applications a consistent way to discover enterprise data and operations. It does not decide what an assistant is authorised to do, how a business operation should fail, or who remains accountable when production state changes.

A mortgage operations analyst asks an internal assistant why an application has not progressed since yesterday. The case record says that a valuation was instructed. The provider callback shows that an appointment was booked. The workflow engine still contains an expired retry task, and the deployment history shows that a callback mapping changed shortly before the case stopped moving.

An assistant with controlled access to those sources can assemble the timeline in seconds. The analyst then asks:

Can you clear the stale task and resend the valuation instruction?

It sounds like a continuation of the investigation. Architecturally, it is a different product.

The investigation reads evidence. The retry changes live state and may contact an external provider. If the request times out, nobody may know whether the provider accepted it. Sending it again could create a duplicate appointment, an unnecessary fee or a case that now requires manual reconciliation.

Both interactions can use the same conversational interface and the same protocol. They cannot rely on the same controls.

This is the enterprise problem behind Model Context Protocol (MCP): how to make operational systems useful to an AI application without giving the model more authority than the person it represents, or bypassing the rules that protect the underlying business process.

The integration problem MCP improves

An enterprise assistant rarely becomes useful by connecting to one system. To explain a stalled case, it may need operational data, integration history, support tickets, deployment records and internal documentation. Without a shared protocol, every AI host and every source system needs a connector with its own discovery model, parameter format and result handling.

MCP standardises that connection. A host can discover the capabilities offered by a server, present relevant context to a model and invoke operations through described contracts. Teams can build clients and servers around one protocol rather than design each host-to-system pairing independently.

MCP is often described as "USB-C for AI applications". The phrase captures one part of the idea: different hosts can discover and use external capabilities through a shared interface instead of bespoke connectors. In an enterprise, however, the shared interface does not make those capabilities equivalent. Reading a policy document, changing verified customer data and rerunning a lending decision may all travel through MCP, but each requires different permissions, safeguards and failure handling.

MCP therefore reduces integration plumbing. The enterprise still has to supply meaning, authority and operational discipline.

What sits behind an MCP connection

MCP uses a host-client-server architecture:

  • The host runs the AI experience, manages model interaction and decides which servers are available in a conversation.
  • An MCP client, created by the host for a particular server, sends protocol requests and returns capabilities and results to the host.
  • The MCP server exposes a selected set of context and operations.
  • Enterprise services behind that server remain responsible for data ownership, business rules and side effects.
flowchart TB
    subgraph HB["Host boundary"]
        U["User"] --> H["AI host"] --> C["MCP client"]
    end
    C -->|"Authenticated MCP request"| S
    subgraph EB["Enterprise boundary"]
        S["MCP server"] --> P["Policy and domain adapter"] --> D["Domain services and systems"]
    end

The network connection between the client and server is an important trust boundary. The host controls which context and tools reach the model, then decides how a proposed call is presented or executed. The enterprise server must make its own decision about whether the caller may perform that operation on that particular business entity.

Server-side capabilities are exposed through three main primitives:

These primitives describe how something is used, not how risky it is. A resource containing a complete customer file may be more sensitive than a tool that creates a draft note. A tool can be a read-only query or a consequential command. The data classification and control model has to be defined separately.

Expose business capabilities, not an API catalogue

It is tempting to generate an MCP server from an existing OpenAPI definition and expose every endpoint as a tool. That produces a quick demonstration. In production, it often gives the model low-level operations, administrative functions and technical parameters that were designed for deterministic callers with detailed system knowledge.

A better server presents a deliberately small capability surface built around business intent:

  • get_case_timeline rather than query_database;
  • create_case_follow_up_task rather than update_workflow_record;
  • retry_valuation_instruction rather than post_provider_payload;
  • request_decision_review rather than set_decision_status.

Each of these tools represents a recognisable business operation, with rules around when it is allowed and what outcome it should produce. retry_valuation_instruction, for example, can check that the current instruction is in a retryable state, apply the same authorisation rules used by other channels and return a result that operations staff understand. Exposing only generic tools such as query_database or post_provider_payload would leave the model to infer those rules from its prompt and assemble the action itself. That is a poor place to put control over a production process.

Schemas should narrow the possible call. Use explicit identifiers, bounded values, format constraints and well-defined output types. Avoid arbitrary SQL, free-form URLs and unbounded request bodies unless the product is intentionally building a privileged engineering tool with a separate risk model.

The MCP server then acts as a controlled adapter. It authenticates the request, selects the capabilities available to that caller, validates the protocol contract and maps the request to a domain command or query. Lending rules, entitlement checks and provider orchestration stay in the existing domain services so that the portal, batch jobs and AI channel receive the same behaviour.

Ownership normally works at two levels. A central platform team can provide transport security, server registration, policy integration, naming standards, observability and compatibility testing. Domain teams should own the tools for valuations, decisioning or case management because they understand the business states and production consequences. A central gateway that automatically converts every internal endpoint into an AI tool tends to accumulate a powerful service account and a catalogue that nobody can reason about safely.

Identity has to reach the system that acts

Authentication at the MCP endpoint is only the first step. For every request, the design should be able to identify the person or workload that initiated it, the host that submitted it and the identity that eventually accessed the system of record.

MCP can connect to a remote server over Streamable HTTP, or the client can launch the server as a local subprocess. The local option is known as the stdio transport: the client writes protocol messages to the process's standard input and reads responses from its standard output. A protected HTTP server can use the OAuth-based authorisation flow defined by MCP. A local subprocess is instead launched and configured by the host, which commonly supplies any credentials it needs through the process configuration or environment. These connection-level arrangements still leave an application question: how will the identity and permissions of the person who requested the action reach the downstream system?

The MCP server may call an internal service using the user's delegated identity, exchange the incoming identity for a downstream token, or use a service identity while enforcing user-level policy itself. Each option can work. What matters is that the initiating user does not disappear behind a shared technical account with broader rights.

In the valuation example, the analyst may be allowed to read the case and its integration timeline but not to issue a new instruction. The server should filter the case data before it reaches the model, then authorise the retry separately against that case and the analyst's current role. A system prompt telling the assistant to respect permissions adds no protection to either step.

Match controls to the consequence

Applying the same confirmation rule to every tool creates two bad outcomes. Either the assistant can do too much without scrutiny, or users face so many approval dialogs that clicking Confirm becomes automatic. Controls should follow business consequence rather than the technical label attached to a tool.

Read versus write is too crude as a classification. A read can disclose an entire portfolio. A small status update may trigger downstream automation. Repeating a provider instruction can create a second real-world activity even if the local database later removes the duplicate record.

Approval becomes meaningful only after the system knows which action is being approved. Before retrying the valuation, the analyst should see the case, provider, current status, reason for the retry and expected effect. Execution should be tied to those values and to the case version on which the decision was made. If the case changes before the command runs, the approval is stale and the operation should stop.

An enterprise-owned step-up flow or a short-lived approval artefact can carry that decision to the executing service. A naked approved: true field cannot. It shows that someone clicked something, but not what they saw, what they agreed to or whether the command changed afterwards.

Treat retrieved content as untrusted data

The assistant investigating a case may read emails, customer documents, issue descriptions, logs and web content. Any of those sources can contain text that looks like an instruction. It may be malicious, copied accidentally from another system or introduced through a compromised integration.

This is indirect prompt injection, also called agent hijacking. An attacker does not need access to the conversation. A sentence embedded in a document can attempt to redirect the assistant, make it disclose data or persuade it to call another tool.

Prompt wording and content filters can reduce obvious attacks, but they are not a dependable enforcement layer. The architecture should limit the effect of a successful injection:

  • expose only the data and tools needed for the user's current task;
  • keep retrieval and command capabilities under separate policies and, where practical, separate credentials;
  • authorise every call independently of the model's explanation;
  • constrain destinations for exports, messages and network requests;
  • test with hostile instructions placed in every content source the assistant may read;
  • monitor unusual tool sequences, repeated denials and unexpected access patterns.

The model can still make a poor decision without an attacker. The same controls protect against ambiguous requests, hallucinated parameters and a user misunderstanding what the assistant is about to do.

Failure semantics belong in the tool contract

The most important behaviour of retry_valuation_instruction may be what happens when no clean answer comes back.

Suppose the integration service sends the instruction and the provider call times out after ten seconds. The provider may never have received it. It may have accepted it while the response was lost. It may also have accepted it and sent a callback that has not yet been processed. A timeout proves uncertainty, not failure.

Automatically calling the tool again is therefore unsafe. The command needs a business idempotency key generated outside the model, a state precondition and a way to query the authoritative outcome. An MCP request identifier is useful for protocol correlation; it is not a substitute for business idempotency across retries, restarts and downstream systems.

A useful tool result makes uncertainty explicit:

{
  "resultType": "complete",
  "content": [
    {
      "type": "text",
      "text": "The provider outcome is not yet known. Do not retry."
    }
  ],
  "structuredContent": {
    "operationId": "valretry-84291",
    "status": "outcome_unknown",
    "caseVersion": 1842,
    "nextAction": "reconcile_with_provider",
    "retryAllowed": false
  }
}

Here, resultType: complete means that the tool call returned a final protocol response. It does not claim that the provider operation succeeded. The host can explain the business uncertainty to the analyst, while operations use the operation identifier and correlation data to trace the command through the MCP server, domain service and provider exchange.

Long-running operations also need an explicit status model. The official MCP Tasks extension can provide an asynchronous task handle and polling where both client and server support it. It does not define whether a valuation can be cancelled, how a duplicate is reconciled or which state is authoritative. Those remain business decisions in the executing service.

Audit the action without copying the whole conversation

A conventional integration log may show that an endpoint returned 202 Accepted. That is not enough when a model proposed the action and a user approved it through a conversational interface.

The audit path should connect the initiating user, host, MCP server and tool version to the target entity, validated command, policy decision, approval and authoritative result. It should also preserve the identifiers required to follow later state changes through downstream services. Where governance requires it, the model and prompt-policy versions can be recorded as part of that chain.

There is a privacy trade-off. Storing the full conversation and every piece of retrieved context can create a second, weakly controlled copy of customer information. In many cases, stable source references, hashes, redacted command snapshots and tiered retention provide better evidence than retaining an indiscriminate transcript.

Observability serves a different purpose. Operations teams need latency, error rates, uncertain outcomes and traces across dependencies. Security teams need denial patterns and anomalous access. Product owners need to know whether the assistant resolves work or simply moves it into another queue. Audit and observability should share correlation identifiers, but they should not be treated as one large log stream.

Build confidence in stages

The safest first release is usually read-oriented and narrow. In this example, the assistant answers one operational question: why has this case stopped progressing? It receives curated case data, integration status and deployment evidence, and it returns source references that the analyst can inspect. The team can measure correctness, access-control behaviour, latency and how much sensitive information enters model context.

The second release can add a reversible operation, such as creating a review task with an explicit owner and reason. That exercises authorisation, idempotency, audit and recovery without allowing the assistant to make a lending decision or contact an external provider.

Only then should the team introduce the valuation retry. By that point, the domain owns the command, the server can enforce case-level permission and state preconditions, and operations have a way to reconcile an uncertain provider outcome. Tests should cover duplicate calls, stale context, permission changes during a conversation, timeouts after a downstream side effect and malicious instructions hidden in retrieved content.

One implementation concern spans all three stages: protocol compatibility. The MCP revision published on 28 July 2026 differs materially from earlier implementations. Hosts and SDKs will not all migrate at the same time, so a production service should pin the revisions it supports, contract-test each host-server combination and give legacy behaviour a defined retirement date.

Back to the stalled valuation

MCP makes it practical for the assistant to gather evidence from several systems and present a coherent case timeline. That is valuable on its own: the analyst gets a faster diagnosis while the underlying case remains unchanged.

The retry requires a larger piece of architecture. The analyst's identity and case permission must survive the full call path. The command must describe a specific business action, check the current case state and carry an idempotency key. Approval must refer to the exact retry the analyst saw. A timeout must produce an honest, traceable outcome rather than an automatic second attempt.

MCP provides a consistent interface for both capabilities. The difference between a useful enterprise assistant and an unsafe production channel is determined behind that interface: in the domain contracts, authority, failure semantics and operational evidence that the implementation puts in place.

Sources and further reading

Related insights

Continue the conversation

More published notes on modernization, integration risk and architecture for critical financial platforms.

Senior architecture discussion

Exploring how AI should reach live enterprise systems?

Quercore Systems helps teams design MCP and AI integration boundaries with the same care they apply to core production architecture.