Skip to main content
Back to all essays
System Design6 min read

Long-running AI agents: resume with current business state

A saved conversation can preserve an old decision. Learn what an agent handoff should retain, what it must refresh and how to test a workflow that pauses overnight.

Overnight agent handoff: save a case checkpoint, refresh the shipment record on resume, compare changes and either continue, revise the proposal or assign an operator.
Open full-size workflow diagram
Read the workflow steps
  1. Persist the case identity, stage, evidence references, proposal version and unresolved work before pausing.
  2. On the next run, load the case and refresh shipment status, rate validity and operator changes from source systems.
  3. Compare the saved proposal with the current records and permission before acting.
  4. Continue a current, permitted proposal; revise one affected by changed facts; route unavailable or ambiguous evidence to its named owner.
Cover graphic for Long-running AI agents: resume with current business state, System Design

An AI assistant prepares a freight exception at the end of a shift. It saves its reasoning and waits for an operator. By morning, the shipment has moved, a carrier rate has expired and someone has corrected the delivery window. The assistant can remember its plan perfectly and still act on the wrong facts. A long-running workflow needs a handoff that distinguishes saved progress from current business state.

The freight case in this article is a proposed design, not a client result or a tested benchmark. It applies lessons from agent engineering to a job that crosses shifts: investigate an exception, prepare a response and wait for a release decision. The question is what the next run needs to establish before that response can proceed.

What the agent engineering research establishes

Anthropic's November 2025 work on long-running coding agents used an initializer to set up the task and subsequent agents to make incremental progress while leaving structured artifacts for the next session. Its experiments found that context compaction alone did not reliably preserve the state needed to finish a web application. Those are findings about coding workflows; the freight handoff below is our proposed operational adaptation. Anthropic: effective harnesses for long-running agents

LangGraph documents checkpoints for thread-scoped graph state and stores for information shared across threads. Its persistence guide also notes that an in-memory checkpointer loses its data when the process restarts. A production handoff therefore needs persistence appropriate to the failure it must survive. LangGraph: persistence and checkpoint scope

Write a case record the next shift can inspect

For our fictional shipment S-318, I would keep a durable case record alongside the integration with the transport system. The record would identify the current stage, the source records used, any prepared proposal and the person responsible for unresolved work. A short model-written summary could help an operator read it. Structured fields would determine which transitions the application permits.

Proposed contents of the case handoff
RecordPurposeCheck on resume
Case and shipment IDsResume the same business jobThe source shipment still exists and belongs to the case
Stage and outstanding workDistinguish prepared, waiting and completed workThe next step remains applicable
Evidence IDs and observed versionsReconstruct the facts behind the proposalSource status, timestamps and versions are current
Proposal version and release decisionKnow exactly what an operator reviewedChanged material facts require a revised decision
Action attempts and recorded outcomesAvoid treating intent as completionReconcile any unresolved downstream result
Owner and review deadlineMake delayed work visible in the queueThe case has an available owner or escalation path

Keep access and retention deliberate. A case may need source references and selected fields without duplicating every email or customer document into agent memory. The operating team should decide what must be retained, who can read it and when it expires. A checkpoint is another data store that needs those decisions.

Walk through the overnight change

At 17:00, the hypothetical assistant reads shipment version 41. It proposes a replacement carrier using a rate that expires at midnight and prepares a customer update based on a 10:00 delivery window. The operator has not released either action. The case enters waiting_for_review with proposal version 3, its evidence references and a named morning owner.

At 08:00, the next run loads that case and retrieves current source records. Shipment version 42 has a 14:00 delivery window, the previous rate has expired and an employee has already assigned a carrier. These changes affect the proposed response. Our application would mark proposal 3 stale and prevent its release. It would prepare a revised summary explaining each change, while preserving the earlier proposal for the audit trail.

Illustrative case state after the morning refresh
case_id: freight-exception-318
shipment_id: S-318
saved_shipment_version: 41
current_shipment_version: 42
previous_proposal: 3
previous_proposal_status: stale
changes: [delivery_window, rate_expiry, carrier_assignment]
next_step: operator_review_of_revised_response
external_action_performed: false

If the source system is unavailable, the saved version remains an observation from yesterday. The screen should show the refresh failure and retain the case for its owner. It should not advance to ready merely because the assistant can write a plausible update. Where the source has no usable version field, the integration needs another comparison contract, such as selected fields plus their observation time, with a documented policy for unresolved changes.

A pause mechanism needs a resume policy

LangGraph's interrupts can pause execution and save graph state while waiting for external input. Resuming uses the same thread ID. The interrupted node restarts from its beginning, so code before the interrupt runs again. Its documentation calls out non-idempotent side effects before a pause as a duplication risk. LangGraph: interrupt and resume semantics

That gives a developer a useful runtime mechanism. The freight application still has to define what a release response means. I would bind the operator's decision to the proposal version and the material facts reviewed, then recheck those facts before execution. A response to proposal 3 would not authorize a revised carrier booking. When the runtime resumes, the application must validate the response rather than treating any returned value as permission.

A checkpoint also cannot establish what a remote system did during an interrupted request. The case record should preserve unresolved attempts so the next run can reconcile them. Our separate retry article explains why a timeout can hide a completed action and why the provider's contract determines whether another attempt is safe. Read: timeouts and safe recovery

Test the handoff across a real restart

For this design, I would ask the team to pause a test case, stop its worker, change the source shipment and resume from a new process. An operator should see the same case, its saved work and the changed facts. Inspect the transport system as well as the agent screen to confirm that the stale proposal caused no external change. This is a suggested verification exercise; we have not run it as a benchmark.

A suggested acceptance review for the handoff
Interruption or changeExpected outcome
Worker stops while waiting for reviewCase, owner and proposal survive in durable storage
Rate expires during the pauseOld quote is invalidated and requires a current alternative
An employee changes the shipmentAssistant exposes the change and preserves the employee's work
Source refresh failsCase stays blocked with an owner and a visible reason
Operator answers an outdated proposalRelease is rejected until the revised proposal is reviewed
Another worker resumes the same caseApplication ownership or concurrency controls prevent conflicting transitions

Before broadening the workflow, measure how often cases resume with stale proposals, how long refresh failures remain unresolved and how much correction the operator performs. Keep completion counts separate from cases still waiting for a decision or confirmation. Start with one exception queue that already has an owner, and have that owner run the morning handoff without an engineer explaining the saved conversation.

Putting this into practice?