I test new work against the thing it replaces and validate experimentation features with the people who run experiments for a living. Four threads where evidence settled the design.
A new drag-and-drop creative builder was meant to replace uploading full ad creatives as flat images and hand-editing CSS to change anything. "Better" was easy to assert; I needed evidence it beat the old way for the people who'd use it.
I ran head-to-head testing with enterprise customers, rating the new builder directly against their current system. The gap was decisive: the new builder scored in the high range, the current system sat near the bottom, and one media customer called it a thousand-percent improvement. Capturing the reasons, like swapping headlines and calls to action without rebuilding an entire creative, turned "better" from an assertion into a documented verdict.
The platform needed a built-in A/B testing flow for marketers, and experimentation UI is unforgiving: traffic logic, locking, and the control case must behave exactly right. The people who'd use it run tests for a living and would catch anything that didn't match real experiments.
I designed the flow against a full set of acceptance criteria, validated it with customers who run experiments daily, and reviewed the whole flow live with a customer ahead of release. Sessions confirmed the core interactions: inline traffic editing, lock behavior, and a split-traffic-equally action. They also caught a subtle issue, a setting applied at the wrong level would have corrupted the metrics, so it went to engineering as a spike before it could ship.
Customers needed a holdout group, a traffic slice that sees nothing, to measure lift; with no real feature, they faked one with invisible variants, causing scroll-lock bugs, CSS conflicts, and support tickets. The feature had to come out of what people were already doing to cope.
I documented the workaround and turned it into requirements for a dedicated holdout: a small default, pixel-based, rendering no creative at all, so measurement stays clean. On whether dismissal settings belong at the campaign or variant level, customer answers split, so it shipped as a configurable default rather than one opinion imposed on everyone.
Before validating with customers, the team had to agree on what it believed. A core system concept rested on unchecked assumptions, and quiet internal disagreement would have turned any customer test into noise.
I ran an internal assumption-mapping session that surfaced more than twenty assumptions and exposed real disagreement on fundamentals, including whether editing a shared template changed existing work. Separating agreement from assumption turned the open questions into an evidence plan, competitor research plus a customer survey through the advisory community, so the team walked into customer validation aligned instead of generating noise.