HealthTech · EMR · Agentic AI

AI Scribe

Designing an agentic documentation partner for physical therapists

Role
Lead Product Designer
Duration
Mar 2025 – present
Company
US HealthTech EMR platform
Year
2025
Platforms
Web, Mobile

The problem

Physical, occupational, and speech therapists spend a large share of their day documenting instead of treating. Every visit needs a structured SOAP note, and for many clinicians that writing happens after hours.

Manual notes also create two risks for the clinic:

  • Inconsistent quality. Style, structure, and completeness vary from therapist to therapist.
  • Compliance exposure. Medicare and other payers require specific documentation, such as Plan of Care recertification windows, KX modifier justification, and co-signature rules. Gaps become claim denials and audit risk.

AI Scribe listens to or takes dictation from a session, transcribes it, and drafts the SOAP note. The therapist’s job shifts from writing the note to reviewing and correcting it. That shift is where the design problem lives: a clinician will only save time if they can trust the draft enough to review it quickly, and correct it without starting over.

My role

I’ve led design for AI Scribe end to end since March 2025, across several phases:

  • Research: competitive teardowns of AI scribe and compliance features across four major PT EMRs, plus usability testing with practicing clinicians before and after launch.
  • Interaction design: the full Review Agent flow, the Scribe Memory framework, the in-app feedback loop, and adoption treatments for new capabilities.
  • Prototyping: from wireframe states to interactive HTML and React prototypes used for stakeholder review and engineering handoff.
  • Cross-platform: adapting the web agent panel into a mobile pattern for the mobile therapist app.
  • Product ownership: writing PRDs with acceptance criteria, phased rollout plans, and risk mitigation.

Engineering owned the transcription and LLM pipeline. Legal handled HIPAA and BAA sign-off. Product leadership and I set adoption targets and priorities together.

How the Review Agent works

  1. The therapist records the session or dictates afterward.
  2. When they stop, the Review Agent analyzes the transcript, extracts clinical details, structures the SOAP sections, and flags missing information. Each step is shown as it happens, so the clinician sees what the AI is doing instead of a blank spinner.
  3. The draft appears next to the transcript. Instead of rewriting by hand, the therapist can instruct the agent (tone, format, emphasis) and have it regenerate sections.
  4. When they’re satisfied, they push the note to SOAP, and the agent session closes. The agent’s job is scoped tightly: from transcript to a note ready to submit, not an open-ended assistant.

Key decisions

Show the work, not a spinner

Early testing made one thing clear: trust in an AI-written clinical note depends as much on process transparency as on output quality. So the agent’s steps (analyzing, extracting, structuring, flagging gaps) are shown as discrete, visible states. Clinicians knew what had been checked before they started reviewing.

Lock the toggle during recording

Letting clinicians switch the Review Agent on or off mid-recording created confusing, error-prone states. During a session the toggle became a read-only status indicator, and the real decision moved to before recording or into settings. We traded a little flexibility for a lot of predictability.

Correct by instruction, not by rewriting

The biggest shift was in how editing works. Rather than rewriting paragraphs, the clinician tells the agent what’s wrong and it regenerates. Editing a note became closer to reviewing a colleague’s draft.

A memory with a clear chain of authority

The first version stored preferences as free-text custom instructions. Clinicians got confused about how clinic rules and personal preferences interacted, and engineering struggled to search, deduplicate, and migrate the text safely.

I redesigned it as Scribe Memory, with three tiers in order of precedence:

  1. Organization rules: set by admins, highest authority, read-only for individual providers
  2. My preferences: each clinician’s own style and phrasing
  3. Patient context: generated automatically for each patient, never mistaken for a standing instruction

This answered a real conflict question: when a clinic rule contradicts a personal preference, the rule wins, visibly. Memory was also organized by specialty (such as MSK, neurological, vestibular, and pelvic health) and by note type (initial evaluation, follow-up, progress, re-evaluation, and discharge).

Let the agent learn from repeated edits

When a clinician makes the same correction over and over, the system notices the pattern and offers to save it as a memory. The scribe gets better the more it’s used, without asking clinicians to write instructions up front.

Cut direct memory editing from the MVP

We originally let clinicians edit memory directly during review. Testing showed they expected the change to affect the current note immediately, but it only applied to future notes. That mismatch would have eroded trust, so we cut it. Instead, the agent proposes saving a preference, and the effect is framed clearly as future-facing.

Simplify memory management

The first memory management UI used overlays on top of overlays and bulk-selection checkboxes. Testing showed they added friction without adding value. We replaced them with inline slide-down forms and single-item actions.

Adoption was a visibility problem

After launch, Review Agent enablement plateaued across clients. The data showed that clinics who found the feature valued it. Many simply hadn’t found it. That reframed the redesign from “make it more valuable” to “make it visible.”

My first idea, a promotional card, broke the panel’s row-based layout and looked out of place once someone had already discovered the feature. So I dropped it for quieter progressive disclosure:

  • an inline accent bar and a “New” badge
  • always-visible feature chips (not hidden behind hover)
  • a permanent collapse to a plain row once the feature is enabled

Closing the quality loop

To learn where the AI fell short, I designed a three-layer feedback system:

  • Inline actions on individual suggestions
  • Panel-level ratings on the draft as a whole
  • A post-push survey covering rating, NPS, accuracy, time saved, and results broken down by specialty

Expanding the value: compliance scoring

Competitive research showed most PT EMRs handle compliance through silent backend checks or manager-only dashboards, with almost no real-time guidance for the treating clinician. Compliance Scoring layers Medicare documentation checks on top of the scribe output to flag gaps before a note is submitted.

One deliberate choice: note-by-note risk scores live in manager-facing views, not as a grade shown to each clinician. The organization gets visibility into audit risk without creating a policing dynamic. We also learned that nudges only change behavior when they’re tied to a concrete consequence, such as a specific CPT code or denial risk, not general “be compliant” messaging.

Taking it mobile

Home-health therapists document between visits, often away from a desk. For the mobile therapist app, I adapted the web right-rail agent panel into a mobile bottom sheet, keeping the same review model while fitting a phone and shorter moments of attention.

What I learned

  • Trust comes from transparency. Clinicians trusted the AI more when they could see what it did, not just what it produced.
  • Low adoption is often a visibility problem. Check discoverability before assuming a feature lacks value.
  • Separate what the org requires from what I prefer. Splitting the two in Memory reduced confusion more than any single UI change.
  • Cut features that break mental models. Dropping direct memory editing made the product more trustworthy, not less capable.