Back to Top

Your AI Agent Is a Distributed System. Design It Like One

Ai agent distributed system

AI agents are moving out of chat demos and into workflows that create tickets, provision infrastructure and modify records. When an agent can write, it inherits every failure mode of distributed systems: timeouts, retries and partial failure. It also adds a new one, which is a caller that can decide by itself to repeat an action.

Tool Calls Are Not Function Calls

Whether you use LangGraph, Spring AI or MCP servers, an agent’s “tool” is usually an HTTP API. Duplicate execution can happen in three ways:

  • Transport retries. The call times out, but the server has already done the work.
  • Orchestrator replays. A crashed workflow resumes from its last checkpoint and runs a step again.
  • Reasoning retries. The model doesn’t see a clear confirmation, so it re-plans and calls the tool again.

Your API gateway can see the first case. It can’t see the third case, because a reasoning retry looks exactly like a new request.

A Real-World Scenario: The Cutover Agent

Consider an agent that runs a cloud migration cutover. It snapshots the database, raises a change request, updates DNS and decommissions the source VM.

The snapshot API times out at 30 seconds but completes at 45, so the agent retries. Two snapshots is an annoyance. Two change requests confuse your CAB. A “decommission VM” step replayed after a checkpoint restore, aimed at the wrong target, is an incident.

Pattern 1: Derive Idempotency Keys From Intent

Generate the key in the orchestrator, never in the LLM. Hash the workflow run ID, the step ID and the business parameters together. If the key came from a random UUID per call, a re-planned call would get a fresh key and slip through.

@PostMapping("/change-requests")
public ResponseEntity<ChangeRequest> create(
        @RequestHeader("Idempotency-Key") String key,
        @RequestBody ChangeRequestDto dto) {

    boolean claimed = redis.opsForValue()
        .setIfAbsent("idem:" + key, "PENDING", Duration.ofHours(48));

    if (!claimed) {
        return ResponseEntity.ok(repository.findByIdempotencyKey(key));
    }
    return ResponseEntity.status(201).body(service.create(dto, key));
}

The atomic setIfAbsent matters. A plain check followed by an insert has a race window under concurrent retries.

Pattern 2: Separate Reads From Writes

Expose read tools freely. Route write tools through a command queue (Kafka or SQS with deduplication) or a transactional outbox. Destructive operations should require human approval in the loop.

Pattern 3: Return State, Not Just Success

A response like "status": "ALREADY_EXISTS", "id": "CHG-1042" tells the model the goal has been achieved. A bare 409 Conflict often triggers another attempt.

Common Mistakes and Trade-offs

  • Key TTLs shorter than your checkpoint resume window. Replays then arrive after the key has expired.
  • Applying idempotency only at the gateway. Downstream side effects such as emails and webhooks still fire twice.
  • Ignoring the cost. An idempotency store adds latency and a new dependency. Budget for it, and monitor it as you would any other part of the critical path.

Key Takeaways

  1. Treat every agent write tool as an at-least-once delivery problem.
  2. Build idempotency keys from workflow and step identity, and generate them outside the model.
  3. Use atomic claim operations, not check-then-insert.
  4. Design tool responses so the model can recognise “already done.”
  5. Gate irreversible actions behind approval, whatever your confidence in the prompt.

 

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Most Popular Posts

Learn React JS with simple Application

Posted on 10 years ago

Bhumi

How to Get IP address in CodeIgniter

Posted on 13 years ago

Bhumi

What is Type Hinting in PHP5

Posted on 14 years ago

Bhumi