State Representation for Agents Navigating Stateful Clinical Applications
Agents need structured state, not screenshots, to navigate clinical software safely.

An agent looking at an EHR screen sees pixels arranged into panels, buttons, and text. A staff member looking at the same screen sees a patient's entire encounter history, knows which tab holds the insurance verification, and remembers that the form will silently reject a date entered in the wrong format. That gap between the two, what is rendered versus what is known, is the subject of this article: how agents build, hold, and update a working model of where they are in a clinical workflow when the interface only ever shows them a fraction of what the software actually contains.
Navigating a clinical application as a state-inference problem
An agent navigating a clinical application never gets to see what the software actually knows. It only gets what the interface chooses to render at that instant, and it has to guess at everything else. A screenshot is a picture of a moment: hidden panels, background processes, document structure, and internal metadata never make it into the frame. The agent's available actions compound the problem, because clicks, keystrokes, and scrolls are coordinate-bound motor movements rather than expressions of intent, so one simple goal, such as "submit this referral", turns into a long chain of precise steps that falls apart the moment a layout shifts or a theme changes. The task an agent is actually performing, then, is to figure out where it is, what has changed since the last look, and what the application still hasn't shown it, using only partial and constantly aging evidence.
Clinical software makes this harder than almost any other domain an agent might work in. EHR sessions routinely run across multiple open tabs, modal dialogs that appear and dismiss themselves on a timer, and payer portal forms where a field doesn't even exist until a prior field has been filled in a particular way. State in these systems doesn't live in one place. A single prior authorization might depend on data entered in the EHR, confirmed in a payer portal, and cross-checked against a claims system, all at once, with no single screen ever showing the full picture. An agent operating in this environment is solving an inference problem, repeatedly, under conditions where the evidence it receives is both incomplete and perishable.
How agents observe a screen
The dominant architecture for computer-use agents takes a screenshot, reasons about it, picks an action, and waits for the next screenshot. That loop misses an entire category of clinical interface behavior: animations, toasts, and transient states that silently decide whether a workflow succeeded or failed. The loop, as described in the Agent-Computer Observation Interfaces Enable Dynamic Computer Use (AOI) paper, works like this: the agent captures a static screenshot, samples an action such as a click, a keystroke, a scroll, a wait, or a signal that the task is done, the interface carries out that action, and only after a buffer period does the next screenshot get taken. Between those two captures, the agent sees nothing at all.
Everything that happens in that interval is invisible to it: an animation playing out, a toast notification that appears and disappears on its own, a loading spinner that signals some background process is running, a spoken prompt from the portal, or a dialog box that opens and closes before the next frame is captured. In a payer portal, a session-timeout warning that flashes once, or a validation error that appears next to a field and vanishes a second later, can quietly derail an entire prior authorization submission, with the agent never registering that anything happened. The same research behind the AOI paper shows that on static benchmarks, agents have gotten substantially better at tasks like form-filling and basic site navigation, but the moment a task involves dynamic or audible content, that progress mostly disappears, and the gap remains large and largely unaddressed. None of this is a hypothetical edge case in clinical software. It's the normal operating condition of a payer portal, where timeouts, validation messages, and audio prompts are routine parts of the form, not rare exceptions the agent might occasionally bump into.
What structured state representation adds over a raw screenshot
One answer is to stop sending agents pictures of software and start sending them the software's actual state. The Agent-Software Interaction Layer, known as ASIL and presented at the Findings of EMNLP 2026, replaces the screenshot-and-click loop with an agent-native interface that exposes software through structured JSON observations and code-executable semantic actions, built on whatever the deepest feasible access path happens to be for a given application. Instead of a flat image, the agent receives an OBSERVATION object organized into specific categories: task metadata, application state, interactive elements, environment context, navigation structure, and a short textual summary of where things stand. Those categories map almost exactly onto the things a screenshot-only agent has to guess at: which fields on a payer portal form are currently clickable, whether a required document has already been attached, what step of a multi-page intake form the session is currently sitting on. Instead of inferring that structure from pixels, the agent is handed it directly.
The difference in outcomes is not incremental. On the ASIL benchmark, the performance gap between agents working from structured state and agents working from screenshots alone is large enough to count as a difference in kind, not a modest improvement at the margins. Other research groups are pursuing the same underlying goal through different technical routes, including vision-based frameworks designed to generalize across platforms, screen-parsing tools such as OmniParser, and methods that augment agents with the accessibility tree, the hidden structural layer operating systems already maintain for screen readers. All of them are trying to surface the same semantic structure that a raw screenshot buries. For a clinical workflow, the practical payoff is straightforward: an agent that knows which fields on a claims appeal form are currently interactive, what the present navigation path through the portal looks like, and what state the application is holding in the background can file that appeal as a sequence of reasoned, semantic steps. It stops being a fragile chain of guesses about where, exactly, on the screen a button happens to sit this week.
Filling the Gaps: Continuous Observation, Keyframes, Audio, and Narration
Structured state solves the problem of what an agent reads at the moment it looks. It does not solve the problem of what happens in between looks, and clinical interfaces generate a steady stream of exactly that kind of event. The AOI paper's answer is the Agent-Computer Observation Interface itself: a perception layer, independent of any particular model, that separates continuous and adaptive observation from the agent's discrete actions, built from three components that only activate when something is actually happening. The first catches visual changes between action steps, things like animations, loading indicators, and dialogs that appear and close on their own. The second transcribes audio only when volume crosses a threshold, catching spoken instructions, portal prompts, and notification chimes that a silent screenshot loop has no channel to receive. The third takes the frames captured by the first component and turns them into persistent text, building a running record the agent can reason over across a long session.
One finding from that research cuts against an intuition it would be easy to hold: the benefit of keyframe capture comes from turning the frames into text that persists, regardless of how cleverly the system chooses which frames to grab. The representation of what was seen determines how much of the session the agent can reason over later, regardless of the sampling strategy used to capture it. The payoff shows up clearly in the audio case: on the spoken-content portion of DynaCU-Bench, agents equipped with AOI solved every task, while screenshot-only agents simply cannot act on audio they never receive in the first place. Clinical relevance follows directly: payer portals with voice-prompted CAPTCHA, session-timeout audio warnings, and notification chimes that signal form-completion states are all in scope for AOI-style observation, and a screenshot-only agent is systematically blind to them. One caveat belongs here for the sake of accuracy: the same research found that on at least one newer model, Gemini 3 Flash, the added keyframe stream actually hurt performance by diluting the image tokens the model had to work with, so these components need to be chosen per model rather than deployed as one fixed setup.
Vast-horizon, repetitive clinical workflows expose state-tracking failures that short tasks hide
A state-tracking mistake that a short task can absorb becomes something far worse once the task repeats hundreds or thousands of times in sequence, because each step's correctness depends on everything that came before it. Prior authorization queues, batches of claims appeals, and stacks of faxed referral orders all share this property. OS-Marathon frames the underlying benchmark problem in three parts: an agent has to read a high-level instruction correctly, carry out the recurring subtasks inside it consistently, and then perform the closing steps once every subtask is actually finished, and current leading agents still struggle with this combination. The workflow shape that matters for healthcare operations is always the same one: a single sub-workflow repeated across a large number of individual cases, where the overall task gets longer simply because the volume of cases grows, not because any one case is harder than the last.
A natural response to that problem is to add an orchestrator that splits the long workflow into separate, smaller subtasks. OS-Marathon's finding on this point runs against the intuition: decomposition alone doesn't fix anything, because errors made by one solver agent accumulate and spread into the next one, since breaking a workflow into pieces doesn't by itself preserve the state that needs to carry over between them. The strategy OS-Marathon explores instead, called GraphDemo, adapts a general-purpose agent to a recurring workflow using a single human demonstration of the sub-workflow's logic. That's the operating principle behind turning an observed staff routine into something an agent can run on its own: show it once, correctly, and let it carry the state of that routine forward across every instance.
Montage Health's referral process shows what's at stake concretely. Faxed referral orders used to depend on one dedicated staff member to transcribe them by hand, and after a referrals-coordinating AI agent was deployed, turnaround times dropped substantially, with a separate Florida-based health system achieving fully automated transcription of thousands of faxed orders without staff intervention, and a far larger volume projected for the following year. Reaching that scale requires the agent to correctly track state across the entire batch of referrals, not just within each individual fax in isolation. The gap between a workflow that collapses under volume and one that scales past it is exactly the gap between isolated subtask execution and a shared state representation carried across the whole batch.
How the access problem limits structured state in practice
Who controls access to the application layer sets the strongest limit on structured state representation in healthcare. Structured observations like the ones ASIL provides depend on instrumentation built into the application itself, and the major EHR vendors, Epic, Oracle Health, and athenahealth among them, control that layer and don't expose it the same way across their products. Standards exist that are meant to solve exactly this problem: FHIR and SNOMED CT were both built to make clinical data portable and semantically consistent across systems. A 2024 scoping review found that full interoperability remains far from solved in practice, citing persistent terminology mapping issues, inconsistent implementations across vendors, and scalability limits, with real deployments often requiring substantial manual mapping and data transformation before the standard actually works as intended.
The computer-use agent model gets around this limitation by refusing to depend on it. Instead of waiting for a vendor to expose structured state through an API, the agent treats the rendered screen itself as the interface, reading it, clicking on it, and typing into it the same way a staff member would, needing nothing more than a login credential. That's the interoperability case for screen-operating agents: they can work across every system already involved in a workflow, EHR, payer portal, third-party scheduling tool alike, without requiring any one of those systems to expose a structured API. A screen-operating agent has to reconstruct state from whatever the interface renders rather than receiving it directly, so the perception and narration mechanisms covered earlier carry the burden of making that reconstruction reliable. They're what make screen-level observation reliable enough to act on.
Two examples show the breadth this approach can reach. An automated prior authorization tool built on a continuously updated library of payer rules automates portal logins and form submissions, including clinical attachments, running the entire workflow through the portal's rendered interface rather than through any vendor-provided API. Myndshft connects to more than 700 payers, covering roughly 93 percent of covered lives for prior authorization, and breadth at that scale is only possible through interface-level access rather than bespoke API deals negotiated payer by payer.
HIPAA compliance for an agent holding state across multi-system workflows
An agent that builds and carries state across an EHR, a payer portal, and a claims system is handling protected health information across systems that were never designed with agent access in mind, and the compliance obligations that follow are a direct consequence of that fact. HIPAA's Minimum Necessary Standard applies at the level of each individual task: the agent's working state should contain only the PHI the current step actually requires, as a limited, task-scoped record rather than everything it has observed over the course of the session. Any vendor whose agent infrastructure touches PHI needs a formal compliance agreement in place, and that requirement applies to the agent platform itself, not only to the EHR or portal it happens to be operating inside.
Audit logging is not optional for AI-assisted action in this environment. End-to-end logs covering every action the agent takes are a compliance requirement, and the narration-to-text mechanism built for continuous observation can double as a natural audit trail when it's structured correctly from the start. Automated payer portal navigation carries a specific version of this exposure: an agent logging into a portal, moving through a multi-step form, and capturing screenshots along the way for audit purposes is, by that very process, handling PHI across systems that were never built to anticipate an agent's presence. SOC 2 Type II certification and HIPAA compliance are the floor any production deployment has to meet before anything else is considered.
The governance infrastructure needed to match this technology hasn't caught up yet. A Q3 2026 healthcare IT market report found that agentic AI in administrative workflows has reached production scale, even as providers remain cautious about letting autonomous agents near tasks that require clinical judgment, and that AI governance across the industry still lags well behind the technology itself. Black Book data cited in that same report found that most healthcare organizations running AI in sustained production had not put complete AI lifecycle controls in place. The capability to track state reliably across a clinical workflow is only half the system. The oversight built to match it is still catching up.


