From APIs to AI Agents in the Supply Chain

Self-healing SC

How connected capabilities can help us recover a customer order across the value chain

A customer expects 100 units on Friday. On Wednesday morning, the warehouse finds that only 60 are available.

Customer service checks the order. A planner looks for stock elsewhere. The warehouse checks what can be picked, and transport looks for a truck. Someone compares the options, gets approval for the extra cost, and makes sure the customer hears a consistent answer.

Much of the information already exists in our applications. The time goes into finding it, understanding the constraints, and coordinating the next step across teams.

This is where I see a practical role for AI agents in supply chain. They can help investigate an exception and coordinate a recovery across the systems we already use. The goal is concrete: the customer receives the complete order on the agreed date, at an acceptable cost.

Let’s follow this order from the first inventory check to confirmed delivery. The figures and API endpoints are illustrative. The code excerpts show the connections, including the process context the agent needs to use them sensibly.

Begin with the value chain

A value chain describes how work creates value across several activities and teams. Here, the journey runs from receiving the order through promising a delivery date, allocating stock, preparing the shipment, delivering it, and issuing the correct invoice.

A warehouse management system, or WMS, manages inventory and warehouse work. A transportation management system, or TMS, manages transport planning and shipment execution. Order management holds the customer order and its promise. Finance manages billing and payment.

The customer experiences the result of their combined work. Stock that arrives after the picking cutoff can still leave the order late. A transport booking is useful only if the warehouse can have the goods ready for collection.

For our case, warehouse A has 60 units available. Warehouse B holds another 50, but must keep 10 to protect other commitments. The customer requires one delivery. The planner can approve up to $500 in extra transport cost.

That gives the recovery process a clear brief: “Find a way to deliver all 100 units by Friday in one shipment, while protecting stock at B and staying within the approved budget.”

APIs make business capabilities available

An API, short for Application Programming Interface, is an agreed way for one application to request information or an action from another.[1]

Think of it as a request form. It specifies what we can ask for, what information we must provide, and what answer to expect. The receiving application still applies its own rules.

For the WMS, we might ask how much stock is available at B. For the TMS, we might ask which transport options could move 40 units from B to A. These requests could look like this, with each response shown underneath:

GET /inventory?site=B&sku=SKU-01
{
  "site": "B", "on_hand": 50,
  "protected": 10, "available_to_transfer": 40
}

GET /quotes?origin=B&destination=A&units=40
{
  "quote_id": "Q-42", "extra_cost": 420, "currency": "USD",
  "arrival_at_A": "Thursday 14:00", "valid_until": "Thursday 10:00"
}

The first answer says that 40 units can be considered for a transfer. The second gives us an option costing $420 extra, with arrival at A expected on Thursday at 14:00. We still need to check that against the warehouse cutoff and confirm the quote before it expires. Actual interfaces should use dated timestamps with time zones and identify the units being moved.

Reading information and changing a business record are different capabilities:

Application Example of reading Example of acting
Order management Read the delivery promise Update an authorized promise date
WMS Check available stock Reserve stock or create a transfer order
TMS Obtain a transport quote Book a shipment
Finance Check invoice status Issue an authorized invoice or credit

For the business, the useful conversation starts with the capability: “We need a reliable way to check stock that is actually available to transfer.” IT can then identify the application, data and API that should support it. Agreeing what “available” means is part of that work.

MCP presents those capabilities to an agent

MCP stands for Model Context Protocol. It provides a standard connection between AI applications and external tools or data.[2]

Imagine a service directory the agent can consult. An entry describes a task, the information it needs and the result it returns. MCP also provides a common way to call that task. Existing APIs often remain underneath it.

For our order, we could expose tools called check_inventory and compare_transport. FastMCP is a Python framework that can turn typed functions into MCP tools. It uses their arguments and descriptions to explain how they should be called.[3]

The connection to the WMS and TMS could be expressed in a small wrapper:

from fastmcp import FastMCP

mcp = FastMCP("Order recovery")

@mcp.tool
async def check_inventory(site: str, sku: str) -> dict:
    """Return stock available to transfer after reservations and protection."""
    response = await wms_api.get(
        "/inventory", params={"site": site, "sku": sku}
    )
    response.raise_for_status()
    return response.json()

@mcp.tool
async def compare_transport(origin: str, destination: str, units: int) -> dict:
    """Return approved carrier quotes for a transfer between sites."""
    response = await tms_api.get("/quotes", params={
        "origin": origin, "destination": destination, "units": units
    })
    response.raise_for_status()
    return response.json()

Here, wms_api and tms_api represent HTTP clients already configured for their respective applications, including authentication and timeouts. The functions call the APIs and pass back their answers. IT publishes this FastMCP service at an authenticated MCP endpoint that the agent can connect to.

The tool descriptions matter. “Available to transfer after reservations and protection” tells the agent more than a vague “get stock” label. The WMS must implement that definition and check it again when stock is reserved.

MCP standardizes the connection. Business rules, data definitions and access permissions still need to be enforced by the services behind it.

Give agents the business process context

The agent also needs a context layer that explains the business process it is working within. I would build it from curated business process documentation: approved process maps, standard operating procedures, decision rules and exception playbooks maintained by the people who own the process.

That documentation should make clear:

  • Which value chain and process apply, and what counts as completion.
  • How the work proceeds, including dependencies and required confirmations.
  • Who can make each decision and when approval is needed.
  • What business terms mean and which local constraints apply.
  • Which capabilities support each step, with examples of their appropriate use.

For our order, the recovery guidance might explain that protected stock is excluded, a transfer reservation comes before booking, and receipt at A must be confirmed before the additional units can be picked. It would also explain how to handle a failed booking and when a change to the customer promise needs review.

The context layer selects the approved guidance for this process and site. It supplies the relevant parts at the start of the case and retrieves further guidance when an exception arises. This use of selected, current information follows the principles of context engineering; loading every company document into every agent would make the relevant guidance harder to find.[7]

The agent combines that guidance with the current MCP tool catalog to choose the capability needed for the next business step. Different WMS or TMS applications can expose tools that serve the same process. The documentation can then be reused across cases and by specialist agents, with each capability mapped to the right step.

Each process document needs an owner, an approved version and a clear scope, such as the sites or customer groups it covers. The case records which guidance was used. WMS and TMS tools supply current operational facts; the context layer supplies the process knowledge needed to interpret them. Together, they help the agent decide what to do next within its permitted authority.

Give the agent a goal and a defined scope

An agent combines a model with instructions and tools. It can choose what to investigate next and use the answers to adjust its approach. A predefined workflow follows a sequence decided in advance.[4]

In this case, the agent might check stock at B, compare transfer options, and investigate another site if the first route cannot meet the deadline. Its analytical work can also use forecasting, simulation or optimization tools when the problem requires them. A language model should ask those tools for quantities and feasible schedules rather than invent them.

Here is the core of a Python agent using the OpenAI Agents SDK. process_context is the relevant guidance selected from the context layer. The case carries the order and its constraints. MCP_URL and MCP_HEADERS configure the connection to the FastMCP service.[5][6]

from agents import Agent, Runner
from agents.mcp import MCPServerStreamableHttp

async def assess_case(case: str, process_context: str):
    async with MCPServerStreamableHttp({
        "url": MCP_URL, "headers": MCP_HEADERS
    }) as tools:
        planner = Agent(
            name="Order recovery planner",
            instructions=(
                "Follow the approved business process guidance below. "
                "Use the tools to check current stock and transport. "
                "Respect protection rules, budget and warehouse cutoffs. "
                "Propose a recovery plan with evidence. "
                "State any missing facts or uncertain assumptions."
                f"\n\nProcess guidance:\n{process_context}"
            ),
            mcp_servers=[tools],
        )
        result = await Runner.run(planner, case, max_turns=8)
        return result.final_output

The SDK handles the loop: the model requests a tool, receives its result, and decides whether another check is needed before returning a proposal.[6] The limit on turns bounds that investigation; a case that cannot be resolved should go to a planner with the facts collected so far.

This service exposes the two reading tools shown above. The planner’s proposal goes through validation and approval before any action is committed. Those controls belong in application code and permissions. An instruction in the agent’s prompt is only guidance.

The roles are easier to see together:

Part Simple reference Role in the example
API A request form Ask the WMS about stock or request a reservation
MCP A common directory and connection for AI tools Present inventory and transport capabilities to the agent
Context layer Curated process guidance Explain the sequence, dependencies and rules for order recovery
Agent A coordinator with a goal and defined authority Investigate how to recover the order
Workflow The recorded sequence of work Carry out approved actions and track their completion

Follow the recovery through to delivery

Once the agent has collected the facts, the decision should be understandable to a planner. For our order, the options might be:

Recovery option Expected customer delivery Extra cost Fit with this order
Wait for replenishment at A Monday $0 Misses Friday
Send 60 units from A and 40 directly from B Friday in two deliveries $260 Customer requires one delivery
Transfer 40 units from B to A then ship all 100 Friday if the transfer meets the cutoff $420 Meets the stated constraints

The transfer could preserve the promise while leaving 10 units at B. Before approval, the process checks warehouse capacity, arrival time, quote validity and any effects on other orders. The proposal should include the alternatives and the reasons for rejecting them.

After the authorized planner approves that specific plan, a workflow reserves stock, creates the transfer and books transport. A request to the WMS might look like this:

POST /transfers
Idempotency-Key: CASE-100-transfer
{
  "case_id": "CASE-100", "order_id": "SO-100", "sku": "SKU-01",
  "source": "B", "destination": "A", "units": 40,
  "approval_id": "APR-42"
}

The approval reference points to a stored decision covering this plan. The service checks the caller’s authority and verifies that the quantity, cost and conditions still match what was approved. Stock availability and the transport quote need to be checked again before committing them.

The idempotency key identifies this particular action. If the API implements that contract, retrying the same request can return the original transfer instead of creating another one. The TMS booking needs its own action identifier and confirmation.

The workflow records the reservation, transfer number and booking reference. If transport cannot be booked, it follows the agreed recovery rule: release the reservation, try an authorized alternative, or escalate. Separate WMS and TMS calls do not form one guaranteed transaction.

Then the physical work takes over. The WMS must confirm receipt at A before the transferred units can be used for picking. The full order can leave when warehouse readiness and outbound transport are confirmed. The case stays open until the agreed customer outcome is observed, such as all 100 units delivered by Friday. Billing follows the applicable rules and delivery evidence.

Make self healing a process we can observe

The analytical part asks what is happening, what is affected and which recovery is feasible. The transactional part reserves stock, books transport or changes an authorized delivery promise. Confirmation connects those decisions to what actually happened.

If the transfer is delayed on Thursday, the recovery case needs another assessment before the outbound cutoff. The system should check whether another option is feasible and identify any new approval needed. A carrier’s revised estimate may be enough to trigger that review.

That is what I mean by a self-healing supply chain: it detects a disruption, assesses recovery options, carries out permitted actions and checks whether the business outcome has recovered. The term is useful only if we can follow that process through to a result.

Some cases will need a person. Conflicting stock records, a missing receipt or an expired quote should produce an escalation that explains what is known and what decision remains open.

Keep decision rights and an audit trail

The business needs to define which actions can run automatically and which require approval. Investigating a shortage might be automatic. Spending $420 on a transfer might require a planner. Changes to customer terms would follow a different authority.

Approval covers the actual proposal. If 40 units become 70, or the cost rises above the approved amount, the system must reassess it. The model’s recommendation should become a structured plan that software can check against those rules.

Each case needs a history showing the facts retrieved and their timestamps, alternatives assessed, rule applied, approval, requested actions and application confirmations. Record the model, tool, instruction and policy versions, plus the process documents and versions used. Protect that history against alteration under the company’s access and retention rules.

One case identifier connects the order, transfer, booking and final outcome. Technical traces help developers investigate failures, while this business history explains what was approved and what happened. A chat transcript alone cannot provide that evidence.

If a connection fails after a reservation request, the reservation may already have succeeded. The workflow should check the original action’s status before repeating it. That matters as much to operations as it does to audit: a duplicate reservation can create the next shortage.

Scale around value chains and shared cases

Several MCP servers can contribute to the same case. An inventory service might expose WMS tools, a transport service might expose TMS tools, and an order service might provide customer constraints. The agent runtime connects to the relevant servers and uses their tools. The business process determines how the work fits together.

I would start with one coordinator. Add specialist agents when their responsibilities or expertise justify them, for example inventory analysis and transport planning. Each specialist should return evidence or a proposal against the same case, and the coordinator should reconcile the result before execution. A fleet needs shared case state and clear ownership of actions. The WMS must reject a reservation if another case has already taken the stock, so two agents cannot both commit the same 40 units.

At higher volumes, case progress must survive a worker restart or a wait for approval. Store it in a durable workflow or case management system, process application events through queues, and limit calls to what the WMS and TMS can handle. A delayed shipment should resume its existing case. Repeated delivery of the same event should not create a second recovery.

The Python excerpt uses the OpenAI Agents SDK for the agent loop. The execution workflow can remain in the company’s existing orchestration platform. Choosing the overall platform should include a practical check: can it resume a case after approval, verify whether a failed call already succeeded, and show the business record across applications?

The same capabilities can support other value chains. Inventory and transport tools could help resolve a supplier delay before production is affected, using the guidance for that process. Returns would add eligibility, inspection and credit capabilities. Each value chain needs a named owner and an agreed definition of completion.

Start with a recurring exception and measure the result

A useful first scope is inventory shortages for one product group across two sites. Curate the recovery procedure with planners, warehouse teams, transport, customer service and IT. Agree on stock definitions, site cutoffs and decision rights, then connect the required capabilities.

First, let the agent collect facts and recommend a plan while people execute the process. Compare the proposals with planner decisions. Next, execute validated plans after approval. Allow a narrow set of cases to run automatically once their constraints and failure paths are understood.

Measure recovered orders, time to confirmed recovery, extra cost and manual effort per case. Track plans rejected by planners, effects on other orders and incorrect or duplicate actions. These measures tell us whether the process is helping the value chain and where it needs correction.

A useful first version can tell a planner: “This order is at risk. Moving 40 units from B could preserve Friday’s delivery for $420, provided the transfer meets Thursday’s cutoff. Here is the evidence and the approval needed.”

Once approved, it follows the case through reservation, movement, warehouse receipt and delivery. If the plan stops working, it brings the next decision back while there is still time to recover the order.

Further reading

  1. IBM, What is an API. A plain introduction to requests between applications.
  2. Model Context Protocol, What is MCP. The standard connection between AI applications and external tools.
  3. FastMCP, Tools. Turning typed Python functions into MCP capabilities.
  4. Anthropic, Building effective agents. The distinction between workflows and agents.
  5. OpenAI, Integrations and observability, and Supply chain copilot with MCP servers. Python SDK connections to MCP services, including Streamable HTTP.
  6. OpenAI, Running agents. How the SDK coordinates the model and its tools.
  7. Anthropic, Effective context engineering for AI agents. Selecting relevant information and retrieving additional context as work proceeds.