Designing trust into AI video understanding

I led product design for VeedoAI, where the real problem wasn't finding moments but trusting them, 0 → 7,000+ users in 4 months
The reframe
Users found things but didn't trust them. I changed it.
The hardest call
I made the AI's limits visible against the category and CEO.
What shipped
Four surfaces. Four months. 7,000+ users. It outlived me.
What I got wrong
I shipped the strategy without instrumenting it.
Project Overview

What is VeedoAI

VeedoAI is an AI video understanding platform for legal, education, and enterprise teams. As the sole designer on a four-person founding team, I designed four product surfaces and the component library behind them, and the product went from zero to 7,000+ users in four months.

The hard part wasn't the interface. It was designing trust into a system that is right most — not all — of the time.

Role
Director of Product & Brand Design (IC, No Reports)
IC · FULL AUTHORITY
Timeline
Feb – May 2025
4MO · 0 TO 7K USERS
Team
CEO, CBO, Director of Growth & Marketing
SMALL, BUT FRIENDLY
Scope
4 surfaces, 3 segments, research, system & brand
COMPREHENSIVE
Results

What the design delivered — and how each number was produced

Every figure on this page carries its provenance. I'd rather show you a smaller honest number than a bigger unattributable one.

0 → 7,000+
Users across five industries, four months post-launch.
Measured in prod
70% faster
To reach a specific moment, vs. participants' own workflow.
Observed in test
1 system, 4 surfaces
The team kept shipping on it after I left, with no designer on staff.
handoff Verified
0
Experiments run on the thesis. I shipped it uninstrumented — the honest gap in this case study.
Not measured
Context & role

In a $7.2B market, design was the only wedge a four-person team had

Every incumbent solved one slice: Twelve Labs had search without collaboration, AnyClip monetization without summarization, Gemini tagging without multimodal interaction. We could not out-model any of them.

twelvelabs.io
AI video search
Gap: No collaboration
anyclip.com
Video monetization
Gap: Limited AI
gemini.google.com
Auto tagging
Gap: No search

That changed what my job was. I wasn't designing screens for an AI product. I was designing the company's differentiation — which meant every decision below is also a decision about where a team with no resources chooses to compete.

I owned
All product design — research through hi-fi and design QA
Brand identity and design language
The research program: 27 professional interviews
The component library and prototyping
I influenced
12-month roadmap sequencing
Feature prioritization against model limits
Positioning and the differentiation narrative
I didn't own
Chukwuma Nwaugha, CEO — model capability vs. interface promises
Ketan Desai, CBO — pricing surface and packaging
Shamal Badhe, Growth — onboarding funnel and activation
The Reframe

The brief was "make video searchable." That was the wrong problem

27 interviews with legal analysts, training managers, and content strategists showed the brief was aimed one layer too shallow. Legal and training participants said the same thing in different words: one wrong summary and they would stop using the tool.

Users weren't failing to find things. They were failing to trust what they found.

So I changed what we were building: not a search tool with good AI, but a comprehension tool that is honest about being probabilistic. Search accuracy wasn't the product problem. Verifiability was.

What it cost the team
The roadmap had collaboration ahead of verification. After the interviews it shipped the other way round: verification affordances went into the MVP and collaboration surfaces moved to a later phase. The reframe cost the team features it had already planned to build.
Whiteboard with sticky notes organizing key insights, time-consuming video review quotes, and design opportunities.
The Evidence Behind It

Three segments, one complaint, three different design answers

27 interviews. Each segment described the same problem in its own vocabulary — and each needed a different thing built.

Teacher in vest and tie helps student writing at desk in classroom with other students seated behind.

🎓 Educators

"Students can't find key lecture moments"
9 interviews
Two people reviewing divorce decree papers across a desk with a Lady Justice statue between them.

⚖️ Legal Teams

"Hours reviewing depositions manually"
7 interviews
Woman adjusting smartphone on ring light for recording or streaming in an indoor setting.

🎥 Creators

"Editing shorts takes longer than filming"
11 interviews
Comparison of broken video issues with needed solutions like intelligent discovery and automated summaries.

The synthesis that moved the brief. Read the right column: every stated need is about acting on what the video contains, not locating it. That asymmetry is what told me the problem was one layer deeper than search.

Framing & success criteria

Nobody could tell me what the problem cost, or what success would look like

Before I could design, I had to build both numbers myself.

The size of the wound

A $7.2B market is not a design brief. Across 27 interviews one figure kept recurring: 3–5 hours per video lost to search and distribution alone. Nobody had multiplied it out. I did.

The model gave the founders a number to price against, told me which segment to design for first, and reframed the product from a convenience into cost recovery. It sizes the wound, not the cure.

A summary that saves you four hours and is wrong once costs more than the four hours it saved.

Retrieval was solved. Verification was not. That gap was the entire opportunity — and it was a design problem, not a model problem, which is the only kind a four-person team could win.

Text explaining the annual cost of a problem per creator and accounts onboarded with model basis stated.
Defining Success

There was no metric for "the AI understood this"

The only proxy anyone used was task completion — did the user find the thing. That measures retrieval, not whether the user believed what they found, which was the actual failure mode. So I modeled comprehension health across three dimensions and proposed them as the product's success criteria.

Dimension 1 · Time
Time to first answer
Not time-in-player. How long from opening a video to having the specific thing you came for. This is the number the 70% improvement refers to.
Dimension 2 · Trust
Verification rate
How often a user clicks from an AI claim through to its source moment. The one I argued mattered most — and the one nobody in the category was instrumenting.
Dimension 3 · Error
Correction rate
How often the AI is wrong in a way the user catches and fixes. Low correction with low verification isn't accuracy — it's users not checking.

Verification rate is not a metric you maximize. Push it to zero and users are trusting blindly — for a legal analyst, a liability. Push it toward one and the AI has saved nobody any time. There's a healthy band in the middle, and where it sits depends on the stakes of the job.

So I wasn't designing to make people trust the AI more. I was designing to make trust proportional to stakes — heavier verification affordances for the legal analyst, lighter for the creator cutting a short. Same model, different trust posture per segment.

The product

Four surfaces, one loop: ask, answer, verify

Ask a question against footage. Get an answer. Verify it against the second it came from.

Ask → Answer → Verify

One interaction connecting search, AI summaries, source moments, and recommendations across four surfaces.

The product

Four surfaces, one loop: ask, answer, verify

Ask a question against footage. Get an answer. Verify it against the second it came from.

Dashboard showing six video project thumbnails with titles, views, time stamps, and editing options.
Surface 01 — Projects

Comprehension starts before you open anything

Every project carries its summary at the list level — the first place "extract, don't watch" shows up, and the reason this isn't a file browser.

Video editing interface showing a makeup tutorial video with AI-generated summary and social platform options.
Surface 02 — Edit

The summary and its source never separate

Recaps generate from the timeline, not a detached transcript — so verification stays one click away at the point where a wrong claim would get baked into a shared clip.

Dashboard showing video analytics with project stats, engagement, watch time, and AI-driven recommendations.
Surface 03 — Analytics

Where attention actually went

The evidence layer under Surface 04. Without it, Refine would be the AI having opinions — the thing the rest of the product argues against.

Dashboard interface for video and blog content optimization with tasks and mark-as-done buttons.
Surface 04 — Refine

Advice carries its evidence too

Every suggestion links back to the behaviour that produced it — the verification rule applied to recommendations. One principle, four surfaces.

Try it yourself

The live prototype

Start at the summary. Click any source chip to jump to the moment it came from. That round trip is the product, everything else on this page is an argument for it.

Open full screen in Figma ↗
How the thinking moved

The wireframes that led to the architecture call

Same screen, same chrome — one structural element changed. The captions record what each version taught, not what it looks like.

Before

Rejected — transcript-first
Wireframe display of a video player interface with search, keywords, topics, OCR, transcript, and scenes panel.
What it taught me
Participants could find words but couldn't act on what they found. Completeness without comprehension — and no amount of polishing the transcript layout fixed that, because the layout wasn't the problem.

After

Chosen — summary-first
Black and white sketch of a laptop screen showing a video player, search, keywords, topics, and transcript sections.
What it taught me
Swapping the content area while keeping every other element pinned proved this was an architecture change, not a redesign. That single observation shaped the final UI — and was the foundation for the before/after below.

↓ These wireframes became Decision 01 below — where the same choice is shown in final production UI, and the cost of making it is named.

The hardest decisions

Three calls, the options I rejected, and what each one cost

Same screen, same chrome — one structural element changed. The captions record what each version taught, not what it looks like.

Before: raw meeting transcript, unstructured wall of text
After: AI-generated summary with decisions and action items
Before — Transcript After — Summary
01 · Information architecture
Summaries first, transcripts one click deeper
✕ Transcript-first
✓ Summary-first
Every competitor defaulted to transcript-first — complete, safe, and not the actual job users were hiring us to do. I bet on comprehension speed over completeness, challenging the entire category’s default.
What it cost
If summaries proved untrustworthy the architecture would have had to be rebuilt. I accepted that risk rather than hedge into a layout that did both badly.
02 · Core interaction
Conversation against footage, not a search bar
✕ Keyword search
✓ In-video AI chat
Keyword search returns timestamps, not answers — we'd have been a worse Twelve Labs. I argued to the CEO that the harder interaction was the moat: conversation was the differentiation, not the feature count.
What it cost
Collaboration surfaces — shared workspaces and commenting — moved to a later phase. Depth instead of coverage.
03 · Failure design
Making the AI's limits visible, not hiding them
✕ Hide uncertainty
✓ Designed fallibility
Where stakes are legal, trust isn’t a feeling — it’s verifiability. I bet admitting fallibility earns more adoption than performing perfection. This applies to any product shipping model output to people accountable for acting on it.
What it cost
A heavier interface, and an argument I kept having. The minimal version tested better on first impression and worse on second use.
The rule I wrote to settle the rest

Progressive AI disclosure — a timing spec, not a principle

AI products fail in one of two directions: they hide what the model can do, or they dump every capability on a first-time user. I rewrote progressive disclosure as what the interface must have proven by which second, and held all four surfaces to it.

Why it mattered beyond the screens: it gave a four-person team a checkable rule for arguments that would otherwise have been taste. "Does this earn its place before the ten-second mark?" settled scope disputes faster than any mockup.

The loop, stage by stage

Every stage had a failure. This is the one I removed from each

Diagram showing a user journey with five steps detailing what broke and what was changed at each stage.
The stage that decided retention
Processing. A user who leaves during the wait never reaches the part that works, which is why the loading sequence got designed before the analytics did.
The stage that had to be earned
Explore. This is where the first AI error lands, and where verification either exists or the user is gone.
Evidence & Limits

How I validated it — and what I couldn't

I ran moderated sessions on UserTesting and Lookback with professionals from all three segments, iterating between rounds. The consistent pattern: users didn't distrust the AI — they distrusted invisible AI. Every fix that worked made the system's reasoning more inspectable.

The change that mattered most came late: participants who skipped first-run setup were the ones who abandoned after their first AI error. Onboarding wasn't polish — it was where trust was established or lost, and I rebuilt it to teach verification rather than tour features.

The honest limit of this evidence
We had no experimentation infrastructure. Nothing here is an A/B test — the 70% figure is an observation from moderated sessions, not a controlled result, and it's labelled that way throughout.
Text outlining an experiment with primary, cells, guardrail, and null hypotheses on verification and trust.
Where the decisions ended up

You don't have to take my word for any of this

Three things I can point at rather than assert. All publicly checkable, none of them mine to edit.

The interaction model persisted
I argued for conversational query over a keyword search bar, using a legal participant's own phrasing. The company's public product listing now describes the feature with the example question "When was the contract mentioned?" — the same question shape, still the primary way the product is explained to buyers.
The system kept shipping
Video Decks launched publicly on 28 June 2025, after my engagement ended and with no designer on staff. It is described as auto-distilling transcripts and visuals into concise slide summaries — the summary-first architecture, applied to a surface I never designed.
The reframe held in the market
An independent buyer-evaluation directory categorises the product's primary job as research insights within video — not video search. That's the reframe I argued for in month one, arriving back as how a third party classifies the company without any input from me.
System & Scope

Built as a system, so one designer could keep pace with four people shipping

The four surfaces share one component library — variants, tokens, handoff-ready. That's the only reason a solo designer kept pace with a founding team's roadmap: a new surface cost composition, not redesign.

The system outlived my engagement. The team kept building on the library after I left, without another designer on staff — the strongest evidence I have that it was built as infrastructure rather than as my own working file.

Desktop — primary
Video review is desk work: long sessions, large timelines, side-by-side reading. I designed the complete loop here first.
Mobile — deferred on purpose
The core loop needs width. Rather than ship a degraded version, I scoped mobile to a later phase and took the reach cost knowingly.
Degraded states — partly debt
AI-is-wrong was designed for. AI-is-slow and processing-failed were not fully specified — I documented them as known debt.
Locale — English-first
Multi-language transcription roadmapped later. Text expansion and non-Latin scripts were known debt I documented rather than solved.
Working With the Founders

Two calls I made, and the pushback I took for them

On a four-person team there is no design org to escalate to. I made these calls, gave the founders the tradeoff, and owned what followed.

Chukwuma Nwaugha
Founder & CEO — AI engineering

The disagreement: he wanted the interface to surface every model capability. I argued it should promise only what the model delivered reliably — an over-promise on a legal summary is unrecoverable, because the user doesn't conclude the feature is weak, they conclude the product lies.

What I did: I held the position and shipped the narrower promise. The cost was a product that demos smaller than it is, in a market where competitors demo everything.

Kit Desai
Co-Founder & CBO — pricing and packaging

The constraint: AI processing is expensive, and a flat price either priced out creators or lost money on heavy users. The tiering question was commercial, but the line had to be drawn somewhere a user could feel — which made it a design problem too.

What I did: I argued the split should fall on capability, not volume — basic search free, advanced AI and faster processing paid — so the free tier still demonstrated the thesis rather than teasing it. A metered free tier would have made the product feel broken before it felt useful.

Shamal Badhe
Director of Growth & Marketing

The disagreement: growth wanted the shortest possible onboarding — every removed screen is a conversion gain. My testing showed the opposite risk: the users who skipped setup were the ones who churned on their first AI error, because nothing had taught them how to verify.

What I did: I traded signup-rate for activation quality and kept the verification moment in the first run, rather than optimizing the number I wasn't accountable for. The cost was friction at the top of the funnel that growth had legitimate reason to dislike.

Outcome

What didn't move, what did, and what changed beyond the product

What did move: four months from first interview to launch, and 7,000+ creators and businesses across five industries — validating the bet that comprehension speed, not feature count, was the differentiator.

The organizational outcome mattered as much: "out-design, not out-model" became how the founding team described the company — in prioritization arguments and against better-funded competitors. The design strategy became the business strategy.

Start with what didn't move
I never measured the thing the whole strategy bet on. The product was designed around verification as the trust mechanism, and we shipped without instrumenting verification rate. Four months in I could tell you adoption was strong, and I could not tell you whether the core thesis was true or whether users were simply trusting blindly — which, for the legal segment, would have been a failure wearing a success's clothes. I made that call on qualitative evidence and accepted the risk knowingly; I'd sequence instrumentation before launch if I ran this again.
Portrait of a middle-aged man with short gray hair wearing large black glasses and a blue checkered shirt.

I had the pleasure of working with Olha during our time together at VeedoAI, where she served as Director of Product and Brand Design. From day one, Olha brought a rare combination of creative vision, design excellence, and strategic thinking that had a transformative impact on our product and brand.

Kit Desai
Co-Founder & CBO
Looking Back

What I'd do differently

01
Instrument the thesis before shipping it

I'd define a behavioural trust metric — verification-click rate against retention — before launch, so the strategy is falsifiable rather than merely plausible.

02
Narrow to the highest-stakes segment first

One core loop across three segments worked, but cost sharpness in each. I'd start with legal, let its verification requirements set the bar, and expand outward — sharpening later is harder than widening later.

03
Pressure-test onboarding in round one

First-run friction surfaced in late-round testing, when changes were expensive. In any 0-to-1, onboarding validation belongs in the first round, not in launch polish.

Questions I get asked about this project

Before you ask in the interview

What was your specific role on VeedoAI?

Sole product designer on a four-person founding team. I owned research through hi-fi and design QA, the brand and design language, the component library, and the 27-interview research program. I influenced roadmap sequencing and positioning. I did not own model architecture, pricing, or paid acquisition.

How were these metrics measured?

The 7,000+ user figure was counted in production over the four months after launch. The 70% figure was observed in moderated sessions against participants' own existing workflow — not a controlled A/B test, and labelled that way. The cost-of-problem figure is modeled, with its inputs shown above.

What would you do differently?

Instrument the thesis before shipping it. The whole strategy rested on verification as the trust mechanism, and verification rate was never instrumented — so the core bet remained unfalsified at launch. I'd sequence instrumentation before launch and narrow to the legal segment first.

Open to new opportunities

Ready To Start?

Book a free 30-minute discovery call

Tell me what you're building.
I'll tell you how I can help and exactly what it will cost.

Currently taking new clients · Typical start: 1–2 weeks from contract

Turning family data into confident action
Dashboard Design
2019 - 2022