Callara Skincare

Callara is an AI skincare startup aimed to help users save time and money buying skincare products that are personalized to their biology. Callara's primary capability is an AI engine that analyzes and builds users' skin profiles. I was hired as a founding product designer. I designed their 0-1 onboarding, core experience, and AI strategy.

Callara app screens — onboarding, AI routine recommendations, product review and skincare goals

Overview

Callara is an AI skincare startup based in Japan with a target market of U.S. consumers, and a forecasted valuation of $2M MRR by next quarter. Callara's founders reached out to me for fractional design help with their prototype to help secure seed funding.

I designed and led the entire AI and UX strategy for Callara's onboarding, resulting in a 67% increase in product input and a 98% increase in users completing skin profile analysis. I delivered a robust prototype and automated a design-to-code pipeline with LLM tools before deploying the live code to the repo.

What I shipped

Five workstreams — from the 0-1 onboarding through a machine-readable design system that kept the AI on spec.

0-1 onboarding — Figma screen flow map

0-1 Onboarding

Here's what we found — scanned items review

Data synthesis and core profile building

Machine-readable semantic design tokens

Machine readable design system

AI rules and strategy — routine score and recommendations

AI rules and strategy

Mobile native — device frame

Mobile native

Find out what your skin needs
Scan your shelf intro
Callara optimizes your routine intro
Login screen
Sign up screen
Scan at least 3 products camera screen
Photo library picker
Here's what we found — scanned items review
Scan results identified
Is this the right product — unable to identify
Is this the right product — manual product entry
Right product confirmed
Skincare goals selection screen
Add products to unlock analysis
Core AI recommendation with 94% match confidence
Core progress screen with 7-day streak

Impact

+0%

Increase in product input. (NS metric)

+0%

Increase in AI recommendation accuracy.

+0%

Increase in time to task completion

TNPS

50% more users recommended Callara's product engine.

Drift

Reducing AI UI styling errors by 80%.

Problem and opportunity space

Challenges

Not clarified MVP and no data funnel

Not clarified MVP and no data funnel

Not clarified MVP and no data funnel

Theinitialprototypewasbasedoffoflargeuserassumptionsthathadnotbeenvalidated.

Ipressedfounderstoaggressivelyvalidate,andremovefeaturesthatweren'tcrucialtotheirnorthstarmetric.

Wedefineduserproductinputasthemainlitmusforproducthealth,andprioritizedfeaturesarounditforphase1rollout.

AI prompted design caused drift

AI prompted design caused drift

AI prompted design caused drift

Callara'steaminitiallyvibecodedtheirprototypeusinggenerativepromptsinFigmamake.Theproblemwasendlessdrifteverytimeasessionwasinitiated.

Theirprototypewasverycrudeandfragmented,withnodistinctbranding,components,hierarchy,userstates,orAIguidelines.

Collaboration

I worked directly with Callara's founders and engineers in live Figma sessions — pairing on the prototype, pressure-testing scope, and aligning the AI behaviour screen by screen.

Live Figma working session with Callara founders and engineers on a video call

Alignment & metric selection

Why prioritize onboarding? An AI system is only as good as the data it receives. I identified weekly product input as our primary metric for retention.

By designing a structured onboarding experience first, we successfully captured critical user data. This immediately improved AI recommendation accuracy, established user trust, and turned early engagement into habitual weekly retention.

North-star metric

Weekly product input — every onboarding decision laddered up to it.

Initial designs, foundations and interaction

Component sheet — button, product card and chip variants extracted from the screens
Scanning products on a bathroom counter — 3 products detectedBrand lifestyle photography — model holding a serum
Is this the right product — AI product verification form with brand, name, type and AM/PM

Setting up skills in Cursor to eliminate drift

Repo

I built a Cursor skill that treats Figma as the single source of truth, then audits the whole repo against it, catching anything that's drifted before it ever ships.

Hover to preview the skill
M+SKILL.md
Preview SKILL.md
~/Documents/GitHub/skin-care/.cursor/skills/sync-design-tokens/SKILL.md

Education, profile, product and guidance

The flow was sequenced to earn trust before asking for data — teach, profile, identify products, then guide.

Education

Education screen

Set expectations up front — a 98.4% accuracy claim primes users to scan their shelf.

Profile

Profile screen

Lifestyle and sensitivity questions build the skin profile that powers every recommendation.

Product

Product screen

AI-identified products with clear review states for anything it could not read.

Guidance

Guidance screen

How to take the perfect photo — guidance that lifts scan accuracy before capture.

Mobile native and Material design constraints

I designed for both platforms natively — respecting iOS Human Interface Guidelines and Android Material Design so the same flow feels at home on each device.

Apple logoiOS — Here's what we found review screen

iOS — Human Interface Guidelines

Material Design logoAndroid Material Design — Here's what we found review screen

Android — Material Design

Accessibility audit

Screen: Lifestyle and sensitivity — a WCAG 2.1 review of colour contrast, touch targets, and semantic structure, then applied to the shipped screen.

Accessibility audit annotations on the Lifestyle and sensitivity screen
Audit — touch targets and focus states flagged
Refined Lifestyle and sensitivity screen after the accessibility pass
Result — 44px targets, AA contrast, clear labels

Data synthesis

Once profiles were built, Callara synthesized product and habit data into a single routine score — the metric users returned for. I explored several ways to surface it and explain the “why”.

AI suitability score — 92 out of 100 with why-this-score reasons and people-like-you
Routine score — Good, could be improved, with product checklist and recommendations
Routine score — Good, could be improved, with a Why this score action