Evidence console

Real job tests, readable proof.

A public-safe UI for the CV-driven InterviewCoach experiments: recent job corpus, mobile interview runs, Spark answer simulations, LangSmith evaluation, and measurement surfaces.

50 recent jobs selected from 91 in-window roles
34 Spark interview reports completed
170 candidate answers submitted in live mobile flow
41 LangSmith eval examples with job URLs

What this proves.

Investors and employers should not need to open raw markdown first. The UI separates product proof, job-fit evidence, and known blockers.

Stable path50/50 deterministic mobile scenarios completed with reports loaded.

Used the same mobile session/report/history API path with private CV held locally and raw answers excluded from public artifacts.

Realistic interviews34 Spark reports across 12 jobs and 12 interview situations.

Five-answer interview loops covered project deep dives, RAG failures, LangGraph support agents, tool safety, coding recovery, and design evolution.

Best-fit signalAverage score 4.34, range 3.98 to 4.65.

Highest current company averages: Trase Systems, Extreme Networks, Bold Business, Pinterest, Elastic, Grafana Labs, and OLX.

Honest boundary7 failed or incomplete Spark scenarios are visible.

Failures are tracked as runtime/stress issues, not hidden. Public reports keep job URLs and sanitized failure causes while excluding raw private answers.

Job-match leaderboard.

These roles came from recent public job pages and were tested against the local CV persona in the mobile interview flow.

Interview situation coverage.

The tests are not one generic chat. They cover common real interview shapes: architecture, coding recovery, RAG incidents, tool safety, and project defense.

8 reportsProject deep dive with business context

40 submitted answers, average score 4.37.

5 reportsAgentic workflow tool-safety review

25 submitted answers, average score 4.38.

3 reportsStreaming/storage design evolution

15 submitted answers, average score 4.46.

3 reportsLangGraph support agent take-home defense

15 submitted answers, average score 4.38.

1 reportDistributed cache HLD/LLD plus TDD

5 submitted answers, score 4.50.

1 reportAI concepts quickfire round

5 submitted answers, score 4.45.

Measurements and review surfaces.

Use these links when reviewing the run. Human-readable UI comes first; JSON and markdown stay available for audit trails.

Source files without making them the UI.

These files remain useful for audit and automation, but the main review path should be the evidence console above.

LeaderboardJob-match leaderboard JSON and readout markdown
Range cohortsLower vs hard cohort JSON and summary markdown
MeasurementMeasurement readout for Grafana, Metabase, Prometheus, and ops visibility.
Initial 50-job run50-job ops summary JSON and progress readout
Privacy boundaryRaw private CV text and raw candidate answers are not published. Public artifacts include job URLs, scores, summaries, gaps, and sanitized runtime failure causes.

Next product work from the tests.

The evidence points to client-journey improvements inside the app, not just more reports.

Evidence checklist before interview

After resume and JD paste, the app should ask for missing metrics, before/after outcomes, architecture scale, and incident examples before starting the interview.

Role-specific question preview

Show the user which interview shapes are likely for the selected job: project deep dive, RAG incident, coding recovery, architecture review, or tool-safety discussion.

Employer/investor demo mode

Expose a polished public-safe demo path that shows coverage, scores, and sources without raw private answers.

Runtime reliability guard

Long Spark scenarios should keep timeout cleanup and retry controls visible so runtime failures are interpreted separately from interview fit.