Designing the human-AI seam for an AI-powered recruitment tool that automates personality assessments while earning trust through explainability and empathy.

Panna is an AI-powered recruitment tool that automates both technical and HR interviews for enterprise clients. I designed the Personality Assessment module, which evaluates candidates' behavioral profiles through a DISC-based questionnaire and automatically generates a data-backed report for recruiters.
The challenge wasn't to "automate hiring," but to create an experience that feels human, unbiased, and interpretable—a space where AI supports, not replaces, recruiter judgment.
Recruiters often rely on subjective impressions during HR screening, resulting in inconsistent and time-intensive decisions. Panna's early prototype automated scoring but failed to communicate how or why AI reached its conclusions—causing skepticism and low adoption.
How might we build an AI-assisted assessment that earns trust and clarifies its reasoning, while still feeling empathetic to candidates?
Scroll to explore the research and design process
Recruiters couldn't validate AI conclusions—no visible rationale.
Candidates found personality questions repetitive and unclear.
The generated report overwhelmed users with raw metrics.
Both parties wanted a human tone and visual clarity in reporting.
Recruiters didn't need smarter AI; they needed smarter communication between AI and user.
Applying lessons from Stanford's Designing AI Products course, I anchored the design around the "human-AI seam"—the transition where people interpret or act on algorithmic outputs.
Show reasoning behind each trait score, not just numbers.
Include confidence levels to visualize model certainty.
Use supportive, neutral language ("growth areas" instead of "weaknesses").
Let recruiters adjust or comment on AI results before final reports.
Mapped current hiring workflows and noted where time and subjectivity created inefficiency. Defined primary KPIs: reduce recruiter effort, increase trust, improve comprehension.
Structured the experience into three flows:
Answering curated paired-choice questions
Model analyzes response patterns + timing consistency
Visualized report with explanations + confidence indicators
Explored three layouts for the results page: card-based, radar chart, and linear report.
Tested with HR users—radar charts won for clarity and familiarity
Simplified the question interface using progress bars and motivational microcopy
Built interactive prototypes in Figma; ran A/B tests comparing color hierarchies (pastel vs. bold DISC colors). Pastel palette was perceived as more trustworthy and calm.
Integrated user feedback into final mockups
Created a design specification document for engineering (font scales, spacing, color states)
Collaborated with the data scientist to ensure explainability phrases aligned with model behavior
Candidate receives an invite link
Completes ~20 scenario-based questions ("Which statement describes you best?")
AI assesses linguistic and behavioral consistency to assign DISC profiles

Assessment questionnaire with paired-choice format

DISC personality assessment results with explainability features
Displays radar visualization of DISC traits
Each trait card (e.g., Steadiness 74% Confidence) links to rationale pop-ups like:
"Responses suggest a calm, supportive communication style. Consistency: 89%."
"Fit Overview" panel breaks down Industry Fit, Role Fit, and Key Qualities
Recruiter can add comments or flag sections for human review
Reports were too technical
Reframed in plain language with interactive tooltips explaining "why."
No indication of reliability
Added Confidence Bars for each trait score.
Recruiters feared bias
Allowed manual annotation before finalizing results.
Candidates disengaged mid-test
Introduced encouraging progress indicators and simple UI states.
Information overload
Used layered visibility — high-level summary first, deeper insights expandable.
Significantly reduced time spent on manual screening
Post-confidence indicators implementation
Via improved test clarity and UX
Up from 61% among new recruiters
Recruiter received a bland "Fit Score: 68%" and couldn't explain why a candidate failed.
Now sees:
Dominance: 72% (High confidence, consistent responses)
Influence: 56% (Medium confidence, variability in tone)
Recommended Role: Account Management, aligns with steady communication pattern
Growth Area: Handles conflict by avoidance; consider live interview follow-up
This shift transformed AI from a judgment tool into a collaborative decision partner.
This project emphasized a lesson that shaped my later AI UX work:
AI adoption doesn't depend on intelligence—it depends on how clearly it speaks to humans.
By focusing on communication, feedback loops, and tone, we designed an assessment that balanced credibility with compassion—a system that augments judgment rather than replacing it.
Introduce voice and facial sentiment analysis for richer context in personality assessments.
Build a transparency dashboard showing model bias and training scope to increase trust.
Develop candidate feedback mode that explains assessment outcomes directly to users.
Designing the 'seam' between AI and humans for predictive restaurant operations
Transforming forest produce supply chain from paper to unified digital ecosystem