← Back to Blog

Designing User Experience for Agentic AI Systems

#agentic-ai#ux#langgraph#human-in-the-loop#production#ai-systems

Updated 2026-09-15 (first published 2026-03-06). The thesis is unchanged, and six months of shipped products and published data now support it. What changed is the mechanics. LangGraph's static interrupt_before is now documented as a debugging tool, not an approval mechanism, so the implementation guidance uses interrupt(). The autonomy slider section now covers classifier-reviewed auto modes and Anthropic's usage data on how oversight changes with experience. New material covers generative UI standards (MCP Apps, A2UI, AG-UI), background agents that deliver pull requests, and mid-run steering. Windsurf is now Devin Desktop.

A team I know built a genuinely impressive research agent. From a single prompt it could search the web, pull documents, synthesize findings, and write a structured report, and it ran on a solid model with working tool integrations and a clean LangGraph graph.

Users hated it.

It was often right, so correctness was not the complaint. They hated it because they had no idea what it was doing: they would submit a task, see a spinner, and wait for thirty seconds or three minutes with no intermediate output or any indication of which step it was on. When it finally returned something, they couldn't tell whether it had actually done what they asked or hallucinated a shortcut, and when it failed, as every production agent does, the error was a generic 500 with no path forward.

The model wasn't the problem. The interface was.

Most agentic AI teams are in this position right now, obsessing over evals, context windows and tool reliability while they ship systems that users don't trust and can't control. Their engineering is sophisticated and their UX is an afterthought.

That is the wrong trade-off.

If you're focused on the engineering side of production readiness, start with 5 Principles for Building Production-Grade Agentic AI Systems and Designing Agentic AI Systems That Survive Production. This article continues from there, at the interface layer.

The Interaction Model Has Changed, But the Interface Hasn't Caught Up

Traditional software follows a simple contract that is deterministic, immediate and predictable: you click a button, and something happens.

mermaid
graph LR
    A[User] --> B[Command]
    B --> C[System]
    C --> D[Result]
    style A fill:#06B6D4,color:#FFFFFF,stroke:#0ea5e9
    style B fill:#EC4899,color:#FFFFFF,stroke:#EC4899
    style C fill:#10B981,color:#FFFFFF,stroke:#10B981
    style D fill:#EAB308,color:#FFFFFF,stroke:#EAB308

Agentic systems break this contract entirely:

mermaid
graph LR
    A[User] --> B[Goal]
    B --> C[Agent Planning]
    C --> D[Tool Selection]
    D --> E[Iteration]
    E --> F[Final Result]
    style A fill:#06B6D4,color:#FFFFFF,stroke:#0ea5e9
    style B fill:#EC4899,color:#FFFFFF,stroke:#EC4899
    style C fill:#10B981,color:#FFFFFF,stroke:#10B981
    style D fill:#EAB308,color:#FFFFFF,stroke:#EAB308
    style E fill:#8B5CF6,color:#FFFFFF,stroke:#8B5CF6
    style F fill:#166534,color:#FFFFFF,stroke:#166534

An agentic system interprets your intent and works out how to get there instead of executing your instruction, so it might take five steps or fifteen, backtrack from a dead end, or make assumptions you didn't authorize.

A better model will not remove this, because it is a fundamental architectural property of autonomous systems. Most teams still ship agentic systems with interfaces designed for the first model: a chat window with a submit button and a loading spinner. That is a command-and-response UI bolted onto a planning and execution engine, and the mismatch is the source of most of the trust failures you see in production.

Part 1 of the Building Real-World Agentic AI Systems with LangGraph guide series covers the conceptual foundation for this shift, from stateless LLM calls to stateful agent execution, in depth.

mermaid
flowchart TD
    U([User]) -->|Goal / Intent| P["Agent Planner\n(LLM)"]

    P -->|Step 1| T1["Tool Call\n(Search / API)"]
    P -->|Step 2| T2["Tool Call\n(Retrieve / DB)"]
    P -->|Step 3| T3["Tool Call\n(Generate / Write)"]

    T1 -->|Result| E["Execution State\n(LangGraph)"]
    T2 -->|Result| E
    T3 -->|Result| E

    E -->|Checkpoint| HiL{"Human-in-the-Loop?"}
    HiL -->|"Yes - interrupt()"| R["User Review\n& Approval"]
    HiL -->|No| OUT["Final Output"]
    R -->|Approved| OUT
    R -->|Edited| P

    style U fill:#4A90E2,color:#fff,stroke:none
    style P fill:#7B68EE,color:#fff,stroke:none
    style T1 fill:#6BCF7F,color:#2C2C2A,stroke:none
    style T2 fill:#6BCF7F,color:#2C2C2A,stroke:none
    style T3 fill:#6BCF7F,color:#2C2C2A,stroke:none
    style E fill:#FFD93D,color:#333,stroke:none
    style HiL fill:#FFA07A,color:#2C2C2A,stroke:none
    style R fill:#E74C3C,color:#fff,stroke:none
    style OUT fill:#98D8C8,color:#333,stroke:none

Every node in this graph is a UX decision point. Where does the user see progress? Where can they intervene? Where does context persist if a step fails? Most teams think only about the LLM and the tools in purple and green, but what users actually experience is the interface layer that wraps all of it.

Users who distrust agents are not being irrational. They are responding rationally to opacity: they can't see what is happening, so they assume the worst.

The Three Questions Every Agentic Interface Must Answer

Before touching a single design decision, get these three questions right for your system:

What is the agent doing right now? Users need to know that the system is working and what it is working on. A spinner does not tell them that, and neither does "Processing...". Show the actual action: Searching documentation for authentication errors, Calling GitHub API to fetch open issues, Generating summary from 4 retrieved documents.

This is functional information. A user who can see the agent searching the wrong source can intervene before it produces garbage output downstream.

Why did it make that decision? Agentic systems choose which tools to call, which information to prioritize, which branches to take. When those choices surface in the UI, users need enough context to evaluate them: less than a full reasoning trace, but enough to say "yes, that makes sense" or "wait, that's not what I meant."

Can I change course? If the answer is no, you have a black box with a submit button instead of a collaborative system. Autonomy without control is a liability.

Interaction Modalities: Match the Interface to the Task

Most production agentic systems default to chat. Text interfaces excel at clarity and traceability, and the rise of AI-augmented terminals like Claude Code and Warp shows how far that modality can stretch when natural language understanding is layered on top.

Text has one core weakness, though: discoverability. Users have no idea what the agent can do unless you tell them. A GUI makes options visible as affordances such as buttons and menus, while a text interface leaves users to figure it out themselves, so they under-use the system, hit invisible capability boundaries, and get frustrated when the agent rejects a request without offering an alternative. Fix it with proactive capability communication: onboarding, contextual suggestions, and graceful redirection when users go out of scope.

Graphical interfaces shine for structured workflows and multi-step processes, and LangSmith, Cursor, and Devin Desktop (Cognition's renamed Windsurf) show what happens when agentic operations get a proper GUI: execution flows become readable, debug cycles shrink, users develop genuine intuition. An emerging pattern here is generative UI, meaning interfaces that create structure dynamically from what the agent produces instead of forcing every output through a predefined template. Coherence is the hard part, because a dynamically generated interface that dumps unstructured information is worse than no interface at all.

In March this pattern had no standard, and it now has three that solve different layers. MCP Apps became the first official Model Context Protocol extension on January 26, 2026, and it works by having a tool declare a ui:// resource that the host renders in a sandboxed iframe inside the conversation. ChatGPT, Claude, Goose and VS Code support it. A2UI from Google (v0.9, April 17, 2026) takes the opposite approach: the agent sends a declarative JSON description, and the client renders it using components from a catalog the client already trusts, which can be your own design system. AG-UI is the event protocol between an agent backend and a frontend, carrying run lifecycle, streamed messages, tool calls and shared state, and A2UI lists it as one of its supported transports.

Choosing between MCP Apps and A2UI is a UX decision disguised as a protocol decision. An MCP App gives a tool author full control of the pixels, so every tool looks like a different product embedded in yours, while A2UI keeps the interface consistent but limits the agent to what your catalog already contains. For agents that act on behalf of users, consistency usually wins, because users learn to read one set of approval cards and status components. A tool author's iframe can also draw a button that looks exactly like your approve button. A catalog component cannot. AG-UI's shared-state sync has production failure modes of its own, covered in The Snapshot Tax: Why AG-UI's STATE_DELTA Drifts in Production.

If your team is hitting the limits of standard React patterns while trying to build these, Frontend Architecture for GenAI: Why Your React Patterns Don't Work Anymore is the right starting point.

Voice works for ambient use cases such as hands-free operations and accessibility, or any scenario where typing is impractical. Its scope is limited by the physics of spoken communication, since speech is slower than reading and dense information delivered aurally can't be skimmed or revisited. Voice agents need to stay conservative.

Match the interface to the level of autonomy: the more consequential the agent's actions, the more structural visibility users need, and a chat bubble is not a sufficient interface for an agent that modifies your database.

The Autonomy Slider: Designing for a Spectrum, Not a Setting

Most teams get one design decision wrong by treating it as binary: how autonomous the agent should be.

Andrej Karpathy framed this well in his June 2025 talk at YC AI Startup School, "Software Is Changing (Again)": effective agentic systems should let users smoothly adjust autonomy across a spectrum, from fully manual control to partial automation to fully autonomous operation. What users need is a slider, not a toggle.

mermaid
flowchart LR
    subgraph MANUAL["🔧 Manual"]
        M1["You do the work\nAgent stays quiet"]
        M2["Tool"]
    end

    subgraph ASSISTED["🤝 Assisted"]
        A1["Agent suggests\nYou approve each step"]
        A2["Copilot"]
    end

    subgraph AUTONOMOUS["🤖 Autonomous"]
        AU1["Agent acts\nYou review after"]
        AU2["Agent"]
    end

    MANUAL -->|"More delegation"| ASSISTED
    ASSISTED -->|"More delegation"| AUTONOMOUS

    style MANUAL fill:#4A90E2,color:#fff,stroke:none
    style ASSISTED fill:#7B68EE,color:#fff,stroke:none
    style AUTONOMOUS fill:#6BCF7F,color:#2C2C2A,stroke:none
    style M1 fill:#3a7bd5,color:#fff,stroke:none
    style M2 fill:#2e6ac4,color:#fff,stroke:none
    style A1 fill:#6a58de,color:#fff,stroke:none
    style A2 fill:#5a48ce,color:#fff,stroke:none
    style AU1 fill:#5bbf6f,color:#2C2C2A,stroke:none
    style AU2 fill:#4aaf5e,color:#2C2C2A,stroke:none

In practice, this maps to three operating modes:

Manual: the agent provides no unsolicited suggestions, so it behaves as a tool that does exactly what you ask and nothing more. This mode is useful when users are doing precision work and need full control, or when they are still building trust with the system.

Ask (Assisted): the agent proactively suggests actions, completions, or next steps, but requires explicit user approval before executing, so the user stays in the decision loop. Cursor suggesting a refactor and waiting for you to accept it is one example, and a code review agent that flags an issue and asks whether to open a PR is another. Throughput stays high and nothing surprises the user.

Agent: the agent executes autonomously within a defined scope and notifies the user of what it did. Users can intervene but don't need to approve each action, which suits well-defined, low-risk, repeatable operations where reviewing every individual step creates more overhead than the risk it mitigates.

These modes aren't static, because user preferences evolve with trust and familiarity. A developer who starts in Manual mode might shift to Agent mode after two weeks, once they have calibrated the system's judgment. Make switching effortless with a visible control instead of a buried setting, and give each mode well-defined, predictable behavior. Nothing erodes trust faster than an agent that behaves inconsistently across modes or surprises users by acting more autonomously than expected.

Beyond being a feature, the autonomy slider is a trust-building mechanism, because giving users control over how much they delegate communicates respect for their expertise and judgment. Agents without that control feel overbearing to some users and underpowered to others, and both groups abandon them.

What the Data Says About the Middle of the Slider

When this article first came out, the middle position meant "the agent asks, the human clicks approve," and usage data published since then shows that this position does not hold up at scale.

Anthropic's own engineering team reported in March 2026 that Claude Code users approve 93% of permission prompts. A prompt that gets approved 93% of the time is not oversight, since it trains people to click without reading, which Anthropic calls approval fatigue. A February 2026 Anthropic study of millions of Claude Code sessions shows what experienced users do instead. New users (under 50 sessions) run with full auto-approve in about 20% of sessions, a share that rises above 40% by 750 sessions, and interrupts rise as well, from 5% of turns at around 10 sessions to about 9% for experienced users. Experienced users keep supervising. They stop approving each action and start watching the run, stepping in when something looks wrong.

Products have moved the slider to match, and the middle position is now a classifier that reviews each action in place of the human. Claude Code's auto mode, generally available since July 2026, runs a two-stage transcript classifier before each tool call. OpenAI's Codex auto-review and the "Automatically approve" setting in Claude in Chrome use the same idea, and two design details in these products matter for UX:

  • A block does not end the run. In Claude Code, a denied action goes back to the agent as a tool result, and the agent looks for a safer path. After 3 denials in a row or 20 in one session, it stops and hands control back to the human. That limit is the handoff from the classifier back to the person.
  • The classifier misses things, and the vendor says so. Anthropic reports a 17% false-negative rate on real overeager actions, and says auto mode does not replace careful human review on high-stakes infrastructure.

So the modern slider has more than three positions, running from Manual (approve each action) through Plan (read and propose, do not change anything) and classifier-reviewed Auto to full bypass inside a sandbox. Plan mode is now a standard mode in Claude Code, Cursor and GitHub Copilot CLI, and all three put it on Shift+Tab. For your own agent the UX lesson is the same, because the positions users actually need are "watch it think", "let it run with a safety net", and "let it run". "Ask me about everything" is mostly a starting point that users leave.

Implementing the Slider in LangGraph

LangGraph's documentation now says static interrupts (interrupt_before and interrupt_after at compile time) are not recommended for human-in-the-loop workflows and describes them as debugging breakpoints. Approval belongs to the interrupt() function: you call it inside a node and resume it with Command(resume=...). Because it runs inside node logic, the decision to pause can depend on the user's autonomy level and on the action itself:

python
from typing import Literal, TypedDictfrom langgraph.types import Command, interruptclass AgentState(TypedDict):    autonomy: Literal["manual", "auto", "full"]    proposed_action: dict    last_result: strdef run_tool(action: dict) -> str:    ...  # your side-effecting tool calldef execute_action(state: AgentState) -> dict:    action = state["proposed_action"]    autonomy = state["autonomy"]    needs_review = autonomy == "manual" or (        autonomy == "auto" and action["irreversible"]    )    if needs_review:        decision = interrupt({"action": action, "autonomy": autonomy})        if decision["type"] == "reject":            return {"last_result": f"rejected: {decision.get('reason', '')}"}        if decision["type"] == "edit":            action = decision["action"]    return {"last_result": run_tool(action)}# Later, when the user clicks Approve in the UI:# graph.invoke(Command(resume={"type": "approve"}), config)

Two details in this sketch come from how interrupt() works, and both break approval UIs when ignored.

First, the autonomy level is read from graph state instead of the per-call config, because LangGraph re-runs the node from the top on resume. If the user changes the autonomy setting while a run is paused, a value read from config changes between the pause and the resume. On the re-run, needs_review evaluates to False, the interrupt never fires, and the action runs even though the user never approved it. Fix the level into state when the run starts.

Second, any side effect placed before the interrupt() call runs again on every resume. Put the tool call after the gate, as the sketch does, or make it idempotent, because otherwise users get duplicate records and notifications for each approval.

If you use LangChain v1's create_agent, HumanInTheLoopMiddleware gives you the same gate per tool, with four decision types: approve, edit, reject and respond. Test the gate itself as well as the UI in front of it. In a LangChain issue from September 2026 (#40200), a conditional when predicate that returns None instead of a boolean causes the interrupt to be skipped, so the tool runs without review.

For a hands-on walkthrough of configuring interrupt policies and building deterministic workflows in LangGraph, see the Building Real-World Agentic AI Systems with LangGraph guide series.

Synchronous vs. Asynchronous: A Design Decision, Not an Implementation Detail

Whether your agent operates synchronously or asynchronously is one of the most consequential UX decisions in agentic system design, and one of the most overlooked. Most teams decide it by what is easier to build, when the decision should follow what the task actually needs.

Synchronous agents operate in real time, with immediate back-and-forth between user and system, as in live chat, voice interfaces or real-time coding assistants. They demand low latency, conversational flow, and context awareness without gaps, and because users expect quick turn-taking, any noticeable pause breaks the interaction rhythm. Design for clarity, brevity, and graceful recovery when the agent misunderstands, which means asking a clarifying question rather than proceeding on a risky assumption.

Asynchronous agents execute tasks in the background and communicate through notifications, summaries, or delivered reports, so users don't wait: they submit work and return to it. Here the design principles flip to persistence, transparency over time, and strong status communication, because users need to know what stage a task is in, when to expect completion, and what happened while they weren't watching. A vague "task completed" notification is nearly as bad as no notification at all.

Teams fall into one failure mode in particular, which is building an asynchronous system with synchronous UI expectations. A long-running agent gets jammed into a chat interface with no status updates, and users watch a spinner for two minutes, wondering whether anything is happening. In the inverse failure, a synchronous agent can't maintain context across the inevitable gaps in a real conversation, so users repeat themselves every few exchanges.

Choose deliberately. If your system starts with a quick synchronous clarification phase and then moves to async background processing, design that handoff explicitly so users know when they have handed control to the agent and what to expect when they return.

Background Agents: The Output Is an Artifact, Not a Message

The biggest async shift since March is in coding agents, and it is a UX pattern other domains can copy. OpenAI's Codex cloud, GitHub Copilot's coding agent, Cursor's cloud agents and Claude Code on the web all run the task away from the user's screen, and none of them is built around a chat reply as the end state. Codex's documented flow is to watch the task logs or let the task run in the background, then review a summary and a diff, then ask for changes or open a pull request. Cursor went one step further in August 2026: cloud agents now subscribe to the pull requests they create and keep working until CI passes and bot review comments are addressed.

Two design choices make this work, and neither depends on code:

  1. The result is a reviewable artifact with its own existing review workflow. A pull request already has a diff view, comments, CI checks and an approve button. The agent reuses a surface where users already know how to evaluate work. For a finance agent, the equivalent is a draft journal entry. For a support agent, it is a queued reply the human can edit.
  2. Status is a small set of named states, shown in one place. OpenAI's Codex app has an activity view that lists chats that are unread, running, or waiting for your response, and an indicator that stays visible while you work in other apps, showing Running, Needs input, Ready or Blocked. Four states are enough. "Needs input" and "Blocked" are the two that justify a notification. "Running" never does.

Asynchronous Processing and Message Queues in Agentic AI Systems covers the infrastructure underneath async agent design: message queues, background processing and state persistence.

Agentic UX in Your Pocket

Execution timelines, editable plans and sync/async handoffs all assume a keyboard, a large screen and focused attention. A phone changes all of those constraints at once, and most agentic UX thinking hasn't caught up to it yet.

Execution timelines, editable plan views and tool activity panels are desktop patterns, and on a 6-inch screen used with one thumb they either become illegible or need so much scrolling that they are unusable. Mobile agentic interfaces need to be radically more opinionated about what to surface and what to collapse. Default to a single-line status indicator such as "Researching your query: 3 of 5 steps complete", with progressive disclosure on tap, instead of a full execution trace.

Input changes too. Voice becomes a first-class interaction mode on mobile in a way it never quite is on desktop, and the reason is the context of use, since the technology is the same. Users on their phone are often between meetings or commuting, away from a desk, where typing a detailed prompt is friction. Short voice commands followed by asynchronous delivery of results form a natural mobile-first pattern, so design for it explicitly instead of defaulting to a text field.

Notification design matters more on mobile than anywhere else, because a push notification is the primary surface for async agent output on a phone, ahead of any dashboard or activity feed. That notification needs enough signal to let the user decide whether to act now or later without opening the app. "Your compliance report is ready: 2 flags require your review before 3pm" is actionable, and "Task complete" is noise.

On mobile, the autonomy slider needs recalibration too, because full autonomous mode is riskier when the user has less oversight infrastructure around them: no second screen to cross-reference and no easy access to full context. Mobile interactions also tend to be glanceable and action-oriented. Consider defaulting to Ask mode on mobile and letting users explicitly opt into Agent mode, rather than inheriting whatever desktop setting they have configured.

Shipping products now do this: in the Claude Code mobile app, you cannot select Bypass permissions for any session. For a Remote Control session, which drives a Claude Code process running on your own machine, the app also removes Auto mode, offering only Manual, Accept edits and Plan. When Remote Control is active, Claude Code sends a push notification to the phone when a long-running task finishes or when Claude needs a decision from you. That applies the notification rule from the next section: interrupt the phone for a decision or a result, never for progress.

Mobile is a different context of use. Attention is fragmented and interaction is ambient, with workflows driven by notifications, so treating a phone as a smaller desktop misses the point. Agentic systems designed only for a 32-inch monitor with full keyboard access will feel broken on a phone even when no feature is missing, because they solve the wrong problem for that context.

Proactive vs. Intrusive: The Hardest Balance in Agent Design

Async agents introduce a problem that synchronous ones don't have: when should the agent reach out to the user, and when should it stay quiet?

Proactivity is genuinely valuable. An agent earns its place when it alerts you to a critical pipeline failure, surfaces a time-sensitive insight before you would have thought to ask, or reminds you of a blocked dependency. Deployed without judgment, the same capability becomes a flood of notifications that trains users to ignore everything the agent sends.

This failure is common and hard to reverse, because once users start dismissing agent notifications reflexively, you have lost the channel, and no amount of "but this one is actually important" recovers it.

Design for context awareness combined with user control. Before it proactively interrupts a user, the agent should be able to answer one question: is this urgent enough to justify interrupting what they are doing right now? During a video meeting, a completed background task delivered by email is fine, and a pop-up alert for the same event is not.

For user control, make notification frequency, delivery channel and escalation thresholds configurable where the interruptions actually happen, so users can tune them in the moment instead of hunting through a settings page. "Don't notify me about this type of event" should be a one-click action on any notification.

Test any proactive behavior by asking whether the notification solves a problem or provides insight the user couldn't have found themselves in a reasonable timeframe. If it does, send it. If it only confirms something they already know, or surfaces information that could easily wait, silence is the better design choice.

Communicating What the Agent Can (and Can't) Do

Users who don't know what the agent is capable of will either dramatically under-use it or hit invisible capability boundaries at the worst possible moment. This failure is easy to overlook during the engineering phase and hard to fix after launch.

Traditional applications solve this with menus, buttons and labels, visual affordances that show available actions so the user does not have to guess. Agentic systems have none of this by default, and text-based agents are the worst off, so capability communication has to be designed in.

A few patterns that work:

Proactive capability introduction. Instead of greeting users with a bare "How can I help?", add one sentence such as "I can help you analyze sales data, draft outreach emails, or debug pipeline failures." That single sentence sets expectations and reduces trial-and-error.

Contextual suggestions. Surface relevant actions based on what the user is currently doing. A user reviewing a document should see "summarize," "extract action items," or "compare with prior version" without having to remember or discover those options on their own.

Graceful out-of-scope handling. When the agent can't do something, it should never simply reject the request. It should redirect instead: "I can't generate invoices directly, but I can draft the line items and hand off to your billing tool." A redirect preserves the relationship and reinforces the agent's utility.

Progressive disclosure. Surface the core five things the agent does well up front, and reveal advanced capabilities as users become more comfortable. Overwhelming new users with the full capability surface is as bad as hiding it entirely.

Aim for legible scope rather than a feature catalogue, so that users can work with the agent confidently and know exactly when to look elsewhere.

Communicating Confidence and Uncertainty

Agentic systems operate on probabilistic outputs, so not every response carries the same degree of certainty, and users need to know the difference, most of all when they are making decisions based on agent output.

Presenting everything with equal confidence is tempting. It feels cleaner, but it is epistemically dishonest, and users eventually find out by acting on something the agent wasn't actually sure about, paying for the mistake, and losing trust in everything the system produces.

Confidence can be communicated in several ways. Explicit statements work well in high-stakes contexts: "Based on the data available, I'm fairly confident in this projection, but the Q4 numbers were incomplete." Power users who want signal density without narrative are better served by visual cues such as subtle color coding or confidence indicators in a graphical interface. Behavioral adjustments are often the most natural: offering a suggestion rather than a firm recommendation when confidence is low, or asking for clarification before committing to an interpretation of ambiguous input.

Calibration matters as much as the mechanism. An agent that expresses false certainty trains users to stop checking, and an agent that hedges everything trains them to ignore it. A confidence signal has to mean something.

When the agent genuinely doesn't know, asking is better than guessing. A focused clarifying question such as "Would you like this for Q3 or the full year?" turns a potential error into a moment of collaboration. Ask one clear question, though, because agents that ask five clarifying questions in a row before doing anything feel bureaucratic and use up user patience fast.

Context Is a UX Problem

Whether your agent remembers what happened before gets treated as a purely technical concern, but it is fundamentally a UX one.

Users experience context loss as the agent being inattentive or obtuse. Having to repeat their project name, their preferences or the constraint they mentioned three messages ago makes the interaction feel transactional and mechanical, and it tells them the system isn't really listening.

Good context retention operates at two levels. Short-term retention holds the details of the current task across the conversation: what has been decided, what the user said they didn't want, and where the workflow currently stands. Long-term retention persists preferences, past patterns, and relevant history across sessions.

Implementation choices shape UX directly. Client-side context is fast but disappears between sessions and devices, while server-side context enables long-term memory but introduces latency and privacy considerations. A hybrid often delivers the best experience, with short-term context kept client-side for responsiveness and long-term context kept server-side for continuity.

Context failures are symmetric. An agent that loses context mid-task forces users to restart from scratch, and an agent that retains too much, or surfaces context the user assumed was ephemeral, feels invasive. Aim for an agent that remembers what helps and forgets what doesn't, which requires explicit thought about what goes into memory and what expires.

When the agent does hit a context gap, it should ask instead of assuming. Graceful recovery with a clear acknowledgment and a targeted clarifying question is much less damaging to trust than a confident answer built on a misremembered premise.

To build persistent state across sessions and across agents, see Building Agents That Remember: State Management in Multi-Agent AI Systems, which covers the implementation side of context retention.

Personalization: The Agent That Learns You

Context retention remembers what happened, and personalization learns from it.

An agent that maintains state knows you mentioned the Q3 constraint three messages ago. An agent that personalizes knows you always prefer bullet summaries over prose, that you work in UK English, that you consistently reject suggestions that involve third-party APIs, and that you tend to start tasks at the architecture level before the implementation.

In practice, personalization takes several forms. With memory of preferences, the agent remembers settings such as notification preferences, output format and verbosity level without you having to restate them: you set them once, and they hold. With style adaptation, the agent adjusts its interaction pattern to observed behavior: if you consistently accept concise responses and skip detailed explanations, it stops offering them. Anticipatory assistance uses past behavior to get ahead of what you will need next, so a project management agent that notices you always ask for a status summary on Monday mornings might start preparing one before you ask.

Overreach is the risk. Personalization done wrong feels invasive, as when an agent references context the user assumed was ephemeral or makes confident assumptions about preferences that haven't actually stabilized yet. Personalization should feel invisible until it is obviously helpful, and when users do notice it, it should feel like attentiveness rather than surveillance.

In practice, that means giving users explicit control to see what the agent has inferred about their preferences, correct it, and reset it. An agent that adapts but can't be corrected will eventually adapt in the wrong direction and stay there.

Putting It Into the Interface

This article opened with three questions every agentic interface must answer, and the autonomy slider, the sync/async split, context retention and confidence signaling are the conceptual foundation for answering them. This section makes that foundation concrete. What the agent is doing right now maps to execution visibility, and why it made a decision maps to plans-before-actions and confidence communication. Whether the user can change course maps to human-in-the-loop interrupts and error recovery.

Expose the Execution, Don't Hide It

Every meaningful action your agent takes should be visible in the primary UI flow instead of buried in a debug log, and a minimal execution timeline looks like this in practice:

mermaid
flowchart TD
    HEADER["🖥️ Agent Execution Timeline: [Stop]"]

    S1["✅ Step 1: Search Documentation\nQuery: 'LangGraph interrupt_before usage'\n→ 4 documents retrieved · 0.8s"]
    S2["✅ Step 2: Call GitHub API\nget_issues(repo='org/repo', state='open')\n→ 12 open issues returned · 1.2s"]
    S3["⏳ Step 3: Generate Summary  [Cancel]\nSynthesizing findings from 16 sources..."]
    S4["○ Step 4: Format and Deliver Output\nPending"]

    HEADER --> S1
    S1 --> S2
    S2 --> S3
    S3 --> S4

    style HEADER fill:#2C3E50,color:#fff,stroke:none
    style S1 fill:#6BCF7F,color:#fff,stroke:none
    style S2 fill:#6BCF7F,color:#fff,stroke:none
    style S3 fill:#FFD93D,color:#333,stroke:none
    style S4 fill:#95A5A6,color:#fff,stroke:none

With this timeline, the user can see what already completed, what is in progress and what is coming next, and they have a cancel button on the active step. That is all. There are no reasoning traces and no debug output, only enough signal to stay oriented and intervene if something looks wrong.

Wanting to hide this because it feels noisy is understandable. Resist it. Users who can see the work trust the output more, even when the output is identical to what they would get from a black box.

Stop Is Not the Only Way to Change Course

Each step in the timeline above has one control, Cancel. That was a reasonable minimum in March, but products now separate three different intents, and a single button forces the user to pick the most destructive one.

ControlWhat the user meansWhen it takes effect
Queue"Do this next, after the current turn"After the agent finishes its current turn
Steer"Keep going, but change direction"At the next safe boundary, usually the next tool call
Stop"Abort this run"Immediately

Steering is the new one. Cursor's August 2026 release changed follow-up messages so they "wait for the next tool call instead of cutting the agent off mid-action." OpenAI's Responses API added mid-turn steering in September 2026 with a response.steer event, and OpenAI's documentation states the limit plainly: steering "does not rewrite output already sent to your application, undo earlier actions, or cancel tools that have already started."

Put that sentence in your UI as well as your API docs. When a user steers, show which already-running actions will still complete, because a user who types "don't email the customer" while the send-email tool is in flight needs to know the email is already going out. Without that, steering produces the worst trust failure an agent can have: the user corrected it, and it did the thing anyway.

Anthropic's usage data shows why this matters. In Claude Code, the most common reason people interrupt is to give missing technical context or corrections, at 32% of interruptions, which makes most interrupts course corrections rather than aborts. Treat Stop as the exception.

Execution visibility for the user and observability for the operator describe the same system from two perspectives. For the operator's side, including what breaks in traditional monitoring when agents go autonomous, see Agentic AI Observability: Why Traditional Monitoring Breaks with Autonomous Systems.

Plans Before Consequential Actions

For any workflow where the agent is about to take an irreversible or high-stakes action, surface the plan first.

text
Proposed plan:1. Pull last 30 days of sales data from the database2. Identify top 5 underperforming SKUs3. Draft email summary to the sales teamApprove / Edit before proceeding?

Good human-in-the-loop design interrupts the agent at the decision boundary before execution begins, instead of interrupting it constantly. In LangGraph, that is an interrupt() call at the top of the node that executes the plan, as in the autonomy slider section above.

Where do you put that boundary? This article's original version said to make approval the default for any action with external side effects, and the data above argues against a rule that broad. In Anthropic's sample of nearly a million tool calls on its public API, only 0.8% of actions appeared to be irreversible. Gate every side effect and you are back to the 93% approval rate and a user who no longer reads the prompt. Gate on irreversibility and scope instead: sending an email to a customer, moving money, deleting data, or touching anything the user did not name in the request. Let a classifier or policy handle the reversible majority, and save the human's attention for the rest.

Plan mode also changes what to do when a run goes wrong. Cursor's documentation recommends that when the agent builds the wrong thing, you revert the changes, refine the plan, and run it again, instead of correcting the run with follow-up prompts. An approved plan is an artifact the user can edit and replay, so build your plan view to make returning to it one click rather than a scroll back through the chat.

Two articles go deeper on the engineering behind this pattern: Multi-Party Authorization: Requiring Human Approval Without Killing Autonomy covers the authorization mechanics, and Consequence Modeling for Agent Systems covers how to predict action impact before execution begins.

Stream Intermediate Output

Waiting for a final result is worse than seeing partial output arrive progressively. Stream tool call results. Stream retrieved document summaries. Stream the plan as it forms. If the final report has five sections, show sections 1 and 2 while 3, 4, and 5 are still generating, so users have something to react to before you are done.

For the implementation layer: Building ChatGPT-Style Streaming in React: FastAPI + Next.js Production Guide covers the full streaming stack end-to-end.

Error Handling Is a UX Problem, Not Just an Engineering Problem

When your agent fails, the question to answer after "did we catch the exception?" is "what does the user do now?"

text
Step 2 failed: GitHub API rate limit exceeded.Options:• Retry in 60 seconds• Skip this data source and continue• Provide a personal access token to increase the limit

Map your failure modes before you build the interface: know what can fail and why, and design a recovery path for each. Ship those paths as first-class UI patterns instead of error handler afterthoughts. In multi-step pipelines, preserve state on failure so the user can pick up from where the agent stopped instead of restarting from scratch.

Teams also consistently get the tone of failure messages wrong, and a cold technical error such as Error 429: rate limit exceeded on resource github_api is accurate and useless to most users. It tells them what broke without telling them what to do, and it communicates nothing about whether the work they submitted is recoverable.

Write agent error messages the way a competent colleague would deliver bad news: acknowledge what happened, be specific about why without jargon, and immediately offer a path forward. "I hit a rate limit on the GitHub API. Your progress is saved, and I can retry automatically in 60 seconds or switch to a cached dataset if you'd rather not wait." That message carries the same information plus context and agency, and the user stays oriented and in control rather than staring at a stack trace.

The Trust Problem Is Structural

One pattern kills agentic system adoption: the agent works correctly 85% of the time, but users can't tell the difference between the 85% and the 15%. Everything looks the same and the UI gives no signal, so users apply equal skepticism to everything the agent produces and conclude that the verification overhead isn't worth the productivity gain.

Performance alone does not build trust in agentic systems. Legibility does, because users need to be able to evaluate outputs as well as receive them.

mermaid
flowchart TD
    OUT["🤖 Agent Output"]

    OUT --> T["Transparency\nSources Shown"]
    OUT --> C["Confidence\nUncertainty Signaled"]

    T --> UE["User Evaluation\nCan verify · Can calibrate"]
    C --> UE

    UE --> TR["✅ Trust"]

    style OUT fill:#4A90E2,color:#fff,stroke:none
    style T fill:#7B68EE,color:#fff,stroke:none
    style C fill:#7B68EE,color:#fff,stroke:none
    style UE fill:#FFD93D,color:#333,stroke:none
    style TR fill:#6BCF7F,color:#fff,stroke:none

In practice, legibility means showing sources when the agent retrieves information, flagging when it is inferring versus when it has direct evidence, making assumptions explicit before acting on them, and signaling uncertainty when it is real. It also means behaving predictably, because users who experience inconsistent agent behavior across similar situations will stop trusting the system entirely, even when it is technically correct.

Trust failures are expensive and slow to recover from, and every negative surprise costs more than ten good interactions can repair. Design defensively: prefer predictable and slightly conservative over capable and occasionally erratic.

The trust problem has a security dimension too. The Agent Trust Problem: Why Security Theater Won't Save Us from Agentic AI covers the architectural patterns that make agents trustworthy from a security standpoint.

What Failure Looks Like: An Anonymized Production Post-Mortem

A fintech team built a document processing agent to automate loan application review. It read uploaded documents, extracted key fields, cross-referenced them against compliance rules, and produced a structured summary for human underwriters. Model accuracy in testing was strong, over 90% on field extraction, so they shipped it.

Three months later, adoption had plateaued at under 30%, and underwriters used the agent only occasionally to generate a first draft before redoing most of the work manually. Assuming the model needed improvement, the team started an expensive re-labeling effort.

A UX audit told a different story: the agent produced clean structured output with no indication of how it had arrived there. When a field was marked "compliant," underwriters had no way to see which document passage had informed that judgment, and when the agent flagged an inconsistency, it showed neither how confident it was nor what it had compared. It occasionally got edge cases wrong, and then there was no recovery path: underwriters had to throw out the entire output and start from scratch, because they couldn't tell which parts to trust and which to re-verify.

Almost all of the fix was UX work. The team added inline source citations that linked every extracted field back to the document region it came from, and a confidence indicator on flagged items that separated high-confidence rule violations from lower-confidence pattern matches. They also introduced a partial edit mode, so underwriters could correct individual fields without invalidating the rest of the output.

Three months after the UX changes, adoption was over 80%, while neither the model nor its accuracy had changed. Only the interface had.

A 90% accurate agent with an opaque interface will lose to a less accurate agent that shows its work. Users do not need perfection. They need enough visibility to apply their own judgment on top of the agent's output, and that collaboration model is the one that holds up in production.

What Good Looks Like: Deep Research Agents

Deep research agents are the clearest production examples of these principles working together: OpenAI's deep research in ChatGPT and Google's Deep Research in Gemini. Both get several things right that most agentic systems get wrong, and the two vendors arrived at nearly the same design.

When you submit a research query, neither one starts generating; each first proposes a research plan. Google's help documentation describes the flow step by step: Gemini creates a plan, you click Edit plan to change it, and research starts only when you click Start research. OpenAI's help center describes the same pattern for ChatGPT: review and modify the proposed plan, and filter which websites the research can use. That is the autonomy slider in action, implemented at the most natural intervention point.

During execution, the user stays in control without having to watch. ChatGPT's deep research shows progress in real time, and OpenAI's help center says you can interrupt the run to refine the focus or change which sources it can access, without restarting. Gemini takes the async route: Google says a report usually takes 5 to 10 minutes, you can leave the chat, and it notifies you when the report is ready, with a push notification on mobile. Both designs make the sync/async handoff from earlier in this article explicit.

Critically, the final reports are sourced, with claims linking back to the documents they came from, which is confidence and uncertainty communication done through structural transparency instead of explicit probability scores. You can check anything without taking the agent's word for it.

Users engage with the output differently as a result: they read it more carefully and trust it more, not because the model is infallible, but because the interface gives them the tools to evaluate it. That is the trust architecture working as intended.

These are not perfect systems. They are slow, and a reviewable plan does not stop a report from favoring breadth over depth. Still, two competing vendors converged on plan review, interruptible runs and cited output, and that convergence is the strongest evidence available that these are the right patterns rather than one product team's taste.

Where This Goes

Eventually, the team from the opening of this article rebuilt their interface without touching the model, the LangGraph graph or the tool integrations. They added an execution timeline, surfaced intermediate results as they arrived, and put a plan confirmation step before the agent started writing. They also wired a LangGraph interrupt to a simple approval screen and rewrote the error messages.

Adoption went from "users hated it" to a tool people actually reached for. The model didn't improve; the interface became legible.

Every production agentic system that has found real adoption repeats this pattern: what wins is the most comprehensible model rather than the most capable one, and the agent users trust enough to let run rather than the one that does the most autonomously.

When this article was first published, the next generation of agentic interfaces was a prediction: collaborative workspaces, plans negotiated before execution begins, generative UI that structures itself around what the task actually requires, and autonomy controls that adapt as trust develops. Six months later, most of that has shipped. Plan mode is on Shift+Tab in the major coding agents, generative UI has open standards, and background agents deliver pull requests instead of chat replies. Classifier-reviewed auto modes let users move along the slider without clicking approve 93% of the time.

Teams that win long-term won't win on interface sophistication either. They will win because they internalized the actual constraint, which is that users don't adopt agents they can't evaluate. Build for legibility first. Capability compounds on top of that, and the reverse order does not work.

It is an interface problem more than a model problem. Start treating it like one.


The most capable agent in the world is useless if users don't trust it enough to act on its outputs.


References


Agentic AI

Follow for more technical deep dives on AI/ML systems, production engineering, and building real-world applications:


Books by Ranjan Kumar

Harness Engineering for Production AI Systems cover

Harness Engineering

The 7 GenAI Architectures cover

The 7 GenAI Architectures

Building Real-World Agentic AI Systems with LangGraph cover

Building Real-World Agentic AI Systems

The ChatML Handbook cover

The ChatML Handbook

The Chat Templates Handbook cover

The Chat Templates Handbook

Comments