Research
Why feedback, why short reps, and what we measure.
Feedback is the active ingredient.
A 2026 randomized study (Louie, Shah, Brunskill, Yang and colleagues, CHI 2026) trained 90+ novice counselors with simulated clients. The group that practiced with structured feedback improved on reflections and questions in a single 75-minute session. The group that practiced without feedback showed no gains, and their measured empathy declined — plausibly because nobody stopped them from sliding into advice. Earlier work (Tanana et al., ClientBot) found immediate feedback on two behaviors produced 91% more reflections. Practice alone is not enough. That’s why every SeatTime rep ends with the coach.
Short, targeted drills beat long sessions.
The deliberate-practice literature (Ericsson; Rousmaniere, Chow and colleagues) is built on brief stimulus–response drills, repeated to fluency, with immediate feedback and rising difficulty. Full 45-minute simulated sessions are where language-model clients drift and students disengage. We build clips first and treat full sessions as “put the corners together,” not the hero.
Students already do this with ChatGPT — without feedback.
A 2025 study at Lamar University found counselor trainees using a general chatbot as a practice client reported lower anxiety and more confidence, but also inconsistency, lack of authenticity, and no structured feedback. We protect the anxiety-reduction benefit (no scores mid-rep, retry always one tap away) and add the part that was missing.
What the coach measures, and where it comes from.
- Counselor Competencies Scale–Revised (CCS-R) — Lambie, Mullen, Swank & Blount, 2015. Part I behaviors map directly to our utterance codes.
- Counseling Skills Scale (CSS) — Eriksen & McAuliffe, 2003.
- MITI 4.2 — Moyers et al. Reflection-to-question ratio, percent complex reflections, percent open questions, MI-adherent vs. non-adherent behavior.
- Barrett-Lennard Relationship Inventory — Empathic understanding subscale, as used in CARE-Bench.
Recent benchmarks (CARE-Bench, PsyCLIENT) use language models as judges scored against these instruments. We do the same, and before launch we calibrate against a human-rated transcript set scored by counselor educators. We’ll publish the agreement figures on this page when we have them.
Calibration results
Coming before launch.
Agreement between the coach and counselor-educator raters on the CCS-R, CSS, and MITI codes, on a shared transcript set. Published here, with the method.
What we deliberately don’t do.
- No persistent clients you “see for months.”
- No real client material, ever. The app blocks obvious identifiers and says so.
- No risk or abuse-disclosure clips without clinical review, scripted client behavior, capped intensity, and a closing note on what real supervision would say.
- No grades.