Skip to content

Research

Why feedback, why short reps, and what we measure.

Feedback is the active ingredient.

A 2026 randomized study (Louie, Shah, Brunskill, Yang and colleagues, CHI 2026) trained 90+ novice counselors with simulated clients. The group that practiced with structured feedback improved on reflections and questions in a single 75-minute session. The group that practiced without feedback showed no gains, and their measured empathy declined — plausibly because nobody stopped them from sliding into advice. Earlier work (Tanana et al., ClientBot) found immediate feedback on two behaviors produced 91% more reflections. Practice alone is not enough. That’s why every SeatTime rep ends with the coach.

Short, targeted drills beat long sessions.

The deliberate-practice literature (Ericsson; Rousmaniere, Chow and colleagues) is built on brief stimulus–response drills, repeated to fluency, with immediate feedback and rising difficulty. Full 45-minute simulated sessions are where language-model clients drift and students disengage. We build clips first and treat full sessions as “put the corners together,” not the hero.

Students already do this with ChatGPT — without feedback.

A 2025 study at Lamar University found counselor trainees using a general chatbot as a practice client reported lower anxiety and more confidence, but also inconsistency, lack of authenticity, and no structured feedback. We protect the anxiety-reduction benefit (no scores mid-rep, retry always one tap away) and add the part that was missing.

What the coach measures, and where it comes from.

  • Counselor Competencies Scale–Revised (CCS-R)Lambie, Mullen, Swank & Blount, 2015. Part I behaviors map directly to our utterance codes.
  • Counseling Skills Scale (CSS)Eriksen & McAuliffe, 2003.
  • MITI 4.2Moyers et al. Reflection-to-question ratio, percent complex reflections, percent open questions, MI-adherent vs. non-adherent behavior.
  • Barrett-Lennard Relationship InventoryEmpathic understanding subscale, as used in CARE-Bench.

Recent benchmarks (CARE-Bench, PsyCLIENT) use language models as judges scored against these instruments. We do the same, and before launch we calibrate against a human-rated transcript set scored by counselor educators. We’ll publish the agreement figures on this page when we have them.

Calibration results

Coming before launch.

Agreement between the coach and counselor-educator raters on the CCS-R, CSS, and MITI codes, on a shared transcript set. Published here, with the method.

What we deliberately don’t do.

  • No persistent clients you “see for months.”
  • No real client material, ever. The app blocks obvious identifiers and says so.
  • No risk or abuse-disclosure clips without clinical review, scripted client behavior, capped intensity, and a closing note on what real supervision would say.
  • No grades.