DESIGNING
FOR WHAT THE
MODEL DOES
NOT KNOW

AI outputs carry uncertainty baked in: the model can be wrong, and users need to see that without losing trust in the product. This case study covers an in-progress, persistent AI agent that follows enterprise users through their journey, and the decisions behind surfacing model outputs honestly, handling uncertainty, and working with ML engineers on what the product can and cannot promise.

Role
Lead Product Designer
Discipline
AI product design, ML collaboration, trust & uncertainty, agentic UX
Context
Enterprise B2B SaaS platform
Threads in this study
03
AGENT USER CONTEXT ACTIONS INSIGHTS AI AGENT A-3
THREAD 01 / 03
Surfacing model outputs
3
Confidence tiers built into the component system
0
Statistical jargon; uncertainty framed in workflow terms
1
Confidence threshold map shared by design and ML
The problem

Enterprise users had to act on model outputs carrying uncertainty a traditional dashboard would hide. A bare number creates false precision that erodes trust when the model is wrong; the full distribution overwhelms, and constant hedging makes the output useless.

What I did

I worked with the ML team to separate what the model knew from what it estimated, then designed a tiered disclosure system: high-confidence outputs surface cleanly, lower-confidence ones carry a visible signal built into the component, with a path to understand why. Mapping the thresholds together kept the UI's confidence language true to what the scores actually meant, so users can tell a solid number from an estimate at a glance.

The goal was calibrated trust: users believe the model when it's right and can push back when it's wrong.

Hide the uncertainty and the product feels broken the first time the model is wrong and users realize they were never told. Design principle, AI output surfaces
SURFACES CLEANLY VISIBLE SIGNAL EXPLAIN PATH WHAT THE MODEL KNOWS THRESHOLDS MAPPED WITH ML ENGINEERS HIGH CONFIDENCE MEDIUM LOW KNOWN, NOT ESTIMATED ESTIMATED UNCERTAIN WHAT THE USER SEES NO FALSE PRECISION, NO CONSTANT HEDGING SURFACES CLEANLY A NUMBER THE USER CAN ACT ON VISIBLE SIGNAL, BUILT IN THE COMPONENT SHOWS ITS UNCERTAINTY A PATH TO UNDERSTAND WHY DISCLOSURE ON DEMAND, NEVER OVERSTATED
Fig. Confidence sets the disclosure: clean when known, signaled when estimated, explained on demand
THREAD 02 / 03
Designing the agent
5+
Journey touchpoints with deliberately designed agent presence
2
Modes: ambient (low footprint), engaged (full context surface)
0
Unconfirmed assumptions surfaced at high-stakes moments
The problem

A one-off AI widget is easy; an agent that follows a user through a whole session, aware of where they came from, what they're trying to do, and what they've already tried, is a different problem entirely. Most enterprise AI surfaces are stateless, so this agent had to hold context across the journey and use it to beat asking from scratch.

What I did

I mapped the journey through the platform to find where an aware agent would cut real friction: transitions between tools, points of failure or confusion, and decisions that sent users elsewhere to look something up. Ambient by default and interruptible on demand, the agent stays quiet until needed and immediately present when it is, so help shows up at the moment of friction instead of becoming one more thing to manage.

Working sessions with product and ML settled what the agent should know versus ask: what the system could infer reliably, what it should confirm, and where assuming too much would feel intrusive. The agent's personality follows those limits, confident where it has signal, honest where it doesn't, which is what makes it worth trusting across a whole session.

An agent that knows where you are but doesn't know what you're trying to do is still just a search box with extra steps. Design constraint, agentic context model
FRICTION POINT THE AGENT ONE AGENT ACROSS THE JOURNEY AMBIENT BY DEFAULT, INTERRUPTIBLE ON DEMAND, QUIET UNTIL NEEDED TOOL A TRANSITION TOOL B POINT OF CONFUSION DECISION WHERE FRICTION LIVED WHERE USERS WENT ELSEWHERE THE AGENT AWARE OF WHERE YOU CAME FROM, WHAT YOU ARE DOING, WHAT YOU HAVE TRIED CONTEXT CARRIED FORWARD, NEVER ASKED FROM SCRATCH
Fig. One agent across the journey, carrying context forward at every step
THREAD 03 / 03
Working with ML constraints
4
Explainability tiers mapped from ML capability to design language
3
Fully designed degraded states: unavailable, low-signal, out-of-distribution
1
Shared ML-design map of product promises and limits
The problem

AI surfaces bring a constraint most product work never sees: the system's behavior isn't fully predictable, and the reasons for its decisions can't always be surfaced. Users want to know why, the model doesn't always have a clean answer, and the product has to hold that gap without making either look bad.

What I did

I ran working sessions with ML engineers to pin down exactly what the model could and couldn't explain about its own outputs, then built a design vocabulary for explainability: what level of explanation each output type allowed, how to communicate it honestly, and how to handle cases where no explanation could be offered.

Graceful degradation covered the rest: the unavailable, underconfident, and out-of-distribution states each got their own copy and recovery paths. A silent model failure is a trust problem; a well-handled one is a product moment.

If the model can't explain it, the product still has to say something. "I don't know" is a valid answer, but it has to be designed. Design principle, explainability systems
FULL PARTIAL NONE WHAT THE MODEL CAN EXPLAIN PINNED DOWN IN WORKING SESSIONS WITH ML ENGINEERS CAN EXPLAIN FULLY A CLEAN REASON EXISTS CAN EXPLAIN PARTLY SIGNALS, NOT A COMPLETE ANSWER CANNOT EXPLAIN THE REASON IS NOT SURFACEABLE THE DESIGN VOCABULARY THE LANGUAGE NEVER CLAIMS MORE THAN THE MODEL KNOWS THE FULL WHY, SHOWN AN HONEST PARTIAL ANSWER SAID PLAINLY: NO EXPLANATION HERE USERS TRUST THE GAP MORE THAN A GUESS
Fig. A graded vocabulary: the language never claims more than the model knows
A-3 / CASE STUDY
← ALL CASE STUDIES