PROVING THE
NEW WAY BEATS
THE OLD ONE,
WITH EVIDENCE

I test new work against the thing it replaces and validate experimentation features with the people who run experiments for a living. Four threads where evidence settled the design.

Role
Lead Product Designer
Discipline
Comparative testing, A/B design, preference research
Context
Enterprise B2B SaaS platform
Threads in this study
04
THREAD 01 / 04
Comparative prototype testing
New builder doubled the old system's usability rating
2
Independent customer sessions converged on the same verdict
1
Clear winner sent to build, evidence attached
The problem

A new drag-and-drop creative builder was meant to replace uploading full ad creatives as flat images and hand-editing CSS to change anything. "Better" was easy to assert; I needed evidence it beat the old way for the people who'd use it.

What I did

I ran head-to-head testing with enterprise customers, rating the new builder directly against their current system. The gap was decisive: the new builder scored in the high range, the current system sat near the bottom, and one media customer called it a thousand-percent improvement. Capturing the reasons, like swapping headlines and calls to action without rebuilding an entire creative, turned "better" from an assertion into a documented verdict.

A thousand-percent improvement on the current system. Enterprise customer, media & publishing
NEW BUILDER CURRENT SYSTEM SAME TASKS, TWO SYSTEMS ENTERPRISE CUSTOMERS RATED BOTH DIRECTLY CURRENT SYSTEM FLAT IMAGES, HAND-EDITED CSS NEW BUILDER DRAG AND DROP, EDIT IN PLACE SWAP A HEADLINE WITHOUT TOUCHING CODE THE VERDICT THE GAP WAS DECISIVE NEW BUILDER SCORED IN THE HIGH RANGE CURRENT SYSTEM NEAR THE BOTTOM A THOUSAND-PERCENT IMPROVEMENT ONE MEDIA CUSTOMER, ON THE RECORD
Fig. Head to head: same tasks on both systems, and a decisive verdict
THREAD 02 / 04
A/B test flow design
17
Acceptance criteria the flow was designed and reviewed against
3
Core interactions validated: inline traffic, lock, split equally
1
Metric-corrupting edge case caught, routed to a spike
The problem

The platform needed a built-in A/B testing flow for marketers, and experimentation UI is unforgiving: traffic logic, locking, and the control case must behave exactly right. The people who'd use it run tests for a living and would catch anything that didn't match real experiments.

What I did

I designed the flow against a full set of acceptance criteria, validated it with customers who run experiments daily, and reviewed the whole flow live with a customer ahead of release. Sessions confirmed the core interactions: inline traffic editing, lock behavior, and a split-traffic-equally action. They also caught a subtle issue, a setting applied at the wrong level would have corrupted the metrics, so it went to engineering as a spike before it could ship.

One variable at a time: creative or settings, never both. Customer experimentation practice, validated in session
CONFIRMED CAUGHT + FIXED DESIGNED TO CRITERIA A FULL SET OF ACCEPTANCE CRITERIA, WRITTEN FIRST THE A/B TEST FLOW TRAFFIC LOGIC, LOCKING, AN UNAMBIGUOUS CONTROL INLINE TRAFFIC EDITING LOCK BEHAVIOR SPLIT TRAFFIC EQUALLY REVIEWED LIVE, BEFORE RELEASE WITH CUSTOMERS WHO RUN TESTS FOR A LIVING LIVE CUSTOMER REVIEW CORE INTERACTIONS CONFIRMED IN SESSION ONE SUBTLE ISSUE CAUGHT, FIXED PRE-RELEASE SHIPPED
Fig. Designed to acceptance criteria, reviewed live, one subtle issue caught before release
THREAD 03 / 04
Holdout group requirements
2/3
Customers preferred campaign-level settings, which set the default
1
Dedicated holdout replacing a bug-prone invisible-variant workaround
0
Rendered creative in the holdout; measurement stays clean
The problem

Customers needed a holdout group, a traffic slice that sees nothing, to measure lift; with no real feature, they faked one with invisible variants, causing scroll-lock bugs, CSS conflicts, and support tickets. The feature had to come out of what people were already doing to cope.

What I did

I documented the workaround and turned it into requirements for a dedicated holdout: a small default, pixel-based, rendering no creative at all, so measurement stays clean. On whether dismissal settings belong at the campaign or variant level, customer answers split, so it shipped as a configurable default rather than one opinion imposed on everyone.

Settings are more precise and aligned with business goals. Customers test creatives more than settings. Preference research, holdout architecture
VARIANT HOLDOUT WORKAROUND BUG THE WORKAROUND CUSTOMERS FAKED IT WITH INVISIBLE VARIANTS INVISIBLE VARIANT A REAL CREATIVE, HIDDEN WITH CSS SCROLL-LOCK BUGS CSS CONFLICTS SUPPORT TICKETS THE FEATURE CAME FROM WHAT PEOPLE ALREADY DID TO COPE TURNED INTO REQUIREMENTS THE DEDICATED HOLDOUT A SLICE THAT RENDERS NOTHING, SO MEASUREMENT STAYS CLEAN VARIANT A 45 VARIANT B 45 10 HOLDOUT: SEES NOTHING SMALL DEFAULT, PIXEL-BASED, NO CREATIVE RENDERED CAMPAIGN LEVEL OR VARIANT LEVEL? CUSTOMERS SPLIT, SO IT SHIPPED AS A CONFIGURABLE DEFAULT LIFT MEASURED AGAINST A CLEAN BASELINE
Fig. The workaround turned into requirements: a holdout slice that renders nothing
THREAD 04 / 04
Assumption mapping
20+
Assumptions surfaced and sorted into agreed versus unverified
3
Concepts with genuine team agreement, isolated from the rest
2
Evidence streams planned to settle open questions
The problem

Before validating with customers, the team had to agree on what it believed. A core system concept rested on unchecked assumptions, and quiet internal disagreement would have turned any customer test into noise.

What I did

I ran an internal assumption-mapping session that surfaced more than twenty assumptions and exposed real disagreement on fundamentals, including whether editing a shared template changed existing work. Separating agreement from assumption turned the open questions into an evidence plan, competitor research plus a customer survey through the advisory community, so the team walked into customer validation aligned instead of generating noise.

Strong agreement on three things, open disagreement on the rest. Better to find out before we sit down with customers. Assumption mapping, validation planning
AGREED DISPUTED EVIDENCE PLAN SURFACE THE ASSUMPTIONS TWENTY-PLUS, MAPPED IN ONE SESSION QUIET DISAGREEMENT, MADE LOUD SORT AGREEMENT VS DISPUTE OPEN QUESTIONS BECOME AN EVIDENCE PLAN GENUINELY AGREED ACTUALLY DISPUTED INCLUDING: DO SHARED TEMPLATE EDITS CHANGE EXISTING WORK? THE EVIDENCE PLAN COMPETITOR RESEARCH PLUS A CUSTOMER STUDY TESTED BEFORE ANY CUSTOMER VALIDATION RAN
Fig. Assumptions surfaced and sorted, open questions turned into an evidence plan
A-5 / CASE STUDY
← ALL CASE STUDIES