CASE STUDY 01 / PERLE_
Cutting ML Project Setup from Hours to Minutes, with AI in the Loop
Setting up a data project on Perle meant hours of manual configuration and a specialist on call. I redesigned it as an AI-assisted flow where the system drafts the setup and the user stays in control.
Company
Perle AI
Role
Product Designer
Focus
AI-Assisted Setup UX
TYPE
Redesign · core flow
Status
● Shipped

tl;dr / 30-second version
The problem
Launching a project required filling ~40 fields across 6 screens in internal vocabulary. Median setup took 3.5 hours, usually with a solutions engineer on a call.
My role
Product designer on the setup flow: research with onboarding calls and config audits, flow architecture, and the review-surface design.
What shipped
A describe → draft → review loop: a plain-language prompt, an AI-drafted configuration with traceable fields, and a paid pilot batch before the full run.
The result
Median time-to-launch 3.5h → 12min, first-attempt completion 38% → 86%, and setup-related support tickets down 61%.
impact
Before anything else, the numbers
3.5h → 12min
Median time to launch a project
38% → 86%
First-attempt completion rate
−61%
Setup-related support tickets
Numbers from 60 days before comparing to 60 days after feature launching
CONTEXT
Everything downstream depends on setup
Perle connects AI teams to expert human data work — labeling, evaluation, and red-teaming for model training. The customers are ML engineers and research leads; the product's job is to turn "we need 10,000 annotated examples" into a running project with the right experts, instructions, and quality gates.
A well-configured project produces clean data; a badly-configured one burns budget and weeks. Setup was also the first thing every new customer touched — it was the first impression.
The brief
Make project setup something a customer can complete alone, on the first try, without losing the rigor the fulfillment side depends on.
THE PROBLEM
Forty fields in a vocabulary nobody spoke
Launching a project required filling ~40 fields across 6 screens: task taxonomy, annotation guidelines, workforce criteria, QA sampling rules, pricing, delivery format. The vocabulary was internal — customers didn't know what a "consensus threshold" was, so they guessed, or they scheduled a call.
Median setup took 3.5 hours spread over days, usually with a solutions engineer on a call.
Only 38% of projects were launched correctly on the first attempt; the rest needed rework after data started coming back wrong.
Sales demos avoided the setup flow entirely — the team knew it didn't sell.
DISCOVERY
The expertise was in the translation
I reviewed 9 onboarding calls and 30 recent project configs.
The pattern: customers could always describe what they wanted in plain language — "rank these two model answers, prefer factual ones, flag unsafe content" — and the solutions engineer would translate that into configuration, live, in about ten minutes.
CORE INSIGHT
The expertise wasn't in filling the form — it was in the translation. And translation is exactly what an LLM does well, as long as a human confirms the result.
design decisions
Four decisions that shaped the flow
Each one traded something. What it cost is part of the decision.
/01
AI drafts, humans own
TRUST
CONFIRMATION
Option A — considered
Auto-launch above 95% confidence
Faster demo, fewer clicks. But a silently wrong project ships thousands of bad items before anyone notices.
Option B — chosen
No auto-launch, ever
Every section is human-confirmed before anything runs, regardless of model confidence.
Why this direction
We traded a slower "wow" for trust, because one silently wrong project costs more than a hundred confirmations.
/02
Keep the form
REVIEW SURFACE
AUDITABILITY
Option A — considered
Replace setup with chat
Conversational only. Elegant for the simple case, hostile to field-level control and procurement review.
Option B — chosen
The form becomes the review surface
Prompt in front, structured configuration behind it — one flow serves both audiences.
Why this direction
Power users and procurement teams still needed field-level control and auditability. Reviewing a plausible draft is faster than composing from zero, and it teaches the vocabulary as a side effect.
/03
Show the why
TRACEABILITY
Every suggestion links to the sentence in the customer's brief that produced it. This cut "is this right?" support questions and made errors easy to spot.
/04
Pilot before scale
RISK
SALES
Sales worried the pilot step would slow deals. The opposite: it became the demo. Prospects saw real data back in hours, which no competitor was showing.
recurring theme
AI earns speed; the human keeps authorship.
Every decision above spends a little friction to keep the customer responsible for what runs.
the product
Describe, review, launch
→ Key screens from the shipped flow. Full prototype available on request.

Describe the task in plain language
The flow opens with a single prompt: "What do you need done?" Customers paste an internal doc, a Slack message, whatever they have. Perle drafts the full configuration from it — taxonomy, guidelines, workforce profile, QA rules.

Review a draft, not a blank form
Every AI-drafted field is visibly marked and traceable — hover shows why the system chose it, linked back to the customer's own words. Editing any field clears the mark. Nothing launches until every section is human-confirmed.

Launch with a safety net
Projects start with a paid pilot batch of 50 items. Results come back within hours; the customer approves the quality or adjusts the config with one round of feedback before the full run. This single mechanic absorbed most of the "wrong config" risk that used to surface after thousands of items.
Your confidence, your control
Define how the AI labels your data
outcomes
What shipped and what it meant
Faster
Setup in minutes
Median time-to-launch dropped from 3.5 hours to 12 minutes, without a call.
Correct
First attempt works
First-attempt completion went from 38% to 86%, and setup-related tickets fell 61%.
Operational
The team scaled
Solutions engineers moved from doing setup to reviewing edge cases — 3× more new projects, no new hires.
business impact
The flow the solutions team used to run by hand is now the product's first impression.
"
"First tool in this space where I didn't need a call to get value on day one."
ML engineer, enterprise customer — onboarding survey
reflections
What designing an AI flow taught me
/01
A draft is a better teacher than a tooltip
Customers learned the platform's vocabulary by correcting a filled-in configuration, not by reading help text next to empty fields.
/02
Traceability is the trust feature
Showing which sentence produced which field did more for confidence in the AI than any accuracy claim could.
/03
The safety net sold the product
The pilot batch was designed to reduce risk and ended up being the strongest thing sales could show.
with hindsight
What I'd push next
PERSONALIZATION
Templates learned per team
Config templates learned from a team's past projects — the draft should get smarter per customer, not just per prompt.
ATTENTION
Confidence-aware review
Spend the user's attention on the 3 fields the model is least sure about, not all 40 equally.
MEASUREMENT
Instrument the draft itself
Tracking which drafted fields get edited most would turn the review surface into a feedback loop for the model.
Next case
Traact Service Center →
