Doc 2 · owner: Rebekah · 2026-08-06 · companion to The Plan

The
Measure

Purpose: produce credible proof for investors (and the launch story) that 3 weeks of Kelas Sekejap improves spoken English. Three instruments that must converge: a blind-rated speaking task, objective transcript metrics, and self-report. We measure change within each student, not absolute level — so the same instrument works for Form 1 and Form 5.

1 · The Speaking Check (in-app, pre + post)

Delivered by the app, not by a person (decided 2026-08-06 — human-run assessments across four sites were a logistical nightmare and don't scale). It's a ~5-minute activity inside the app: the pre runs as the student's very first activity — Day 1 stays locked until it's done, which guarantees a true baseline and enforces "joining = doing both" — and the post unlocks after Day 21 as the final activity.

Disclosed upfront as a joining condition (teacher pitch, parent letter, kickoff commitment). In everything student-facing it is a "short speaking activity — not an exam", never a "test": exam framing scares off exactly the weaker students we recruited, and the friendly framing must be identical in both rounds so the recordings are comparable.

Flow per round (~5 min, identical pre and post)

  1. Intro screen (B1 copy): "This is not an exam. There are no wrong answers. Find a quiet place. Speak in English and try to keep talking until the timer ends."
  2. Confidence check (one tap): "How confident do you feel speaking English?" — 1–10 scale. Asked identically pre and post; this is the before/after confidence measure, captured in-app so collection is 100%.
  3. Warm-up (not recorded): one easy question ("What did you do last weekend?"), ~30 seconds — settles nerves so Task 1 isn't cold.
  4. Task 1 — Describe (2 min, recorded): personal-descriptive prompt with 3 guide points on screen, countdown timer visible.
  5. Task 2 — Opinion (1 min, recorded): simple agree/disagree question.
  6. Confirmation screen: upload confirmed ("We got it — thank you!"). A failed upload retries until confirmed and flags the account for follow-up.

Rules that protect the measure

Unsupervised-conditions caveat: nobody is watching, so a student could sit in noise or get help. Mitigations: kickoff instruction is "do it tonight, in a quiet room, alone"; timestamps recorded; Eugene/teachers chase stragglers on WhatsApp. The part that matters for validity: conditions are identically unsupervised in both rounds, so the change score stays meaningful.

Task versions (counterbalanced)

Half the students get A pre / B post; the other half B pre / A post — assigned automatically. It kills the "they just saw the question before" objection and cancels out one version being slightly easier.

Version A

Describe: "Tell me about a person who is important to you." Guide points: Who is this person? · What do you do together? · Why are they important to you?

Opinion: "Some people say all students must wear school uniforms. Do you agree? Why?"

Version B

Describe: "Tell me about a place you like to go." Guide points: Where is it? · What do you do there? · Why do you like it?

Opinion: "Some people say students should be allowed to use phones in school. Do you agree? Why?"

2 · Rubric (three transcript-gaugeable criteria)

Three criteria, 1–5 each, total /15. Pronunciation and delivery were deliberately dropped: they can only be judged from audio, and the scoring runs on transcripts so a model can score every recording consistently. Confidence still shows up in the self-report and the objective metrics; pronunciation is out of scope for this study.

Criterion135
FluencyMostly silence; gives up; isolated wordsFrequent pauses but recovers; short connected runsKeeps talking for the full time; speech flows
VocabularyA few basic words, heavy repetitionEveryday words, some repetition, occasional wrong wordVaried word choice; finds a way around gaps
GrammarErrors block understandingErrors frequent but meaning survives; simple sentencesMostly accurate; attempts longer sentences

Scoring protocol — model-scored, human-validated

  1. Pull the transcripts and rename each to a random code (R014); the code→student/round key lives in one spreadsheet only the admin (Rebekah) sees. Strip anything that reveals the round.
  2. Claude scores every transcript against the rubric — one batch run, fixed JSON output schema (score + one-line justification per criterion). Blind by construction: the model is never told whether a transcript is pre or post. Perfectly consistent, rerunnable, no rater fatigue. Cost for the whole study: under RM5.
  3. Rebekah blind-rates a validation sample — ~20 recordings (audio + transcript), same three criteria, shuffled, round unknown. About 1.5 hours.
  4. Report model–human agreement on the sample (% of criterion scores within 1 point) — the line that pre-empts the "AI graded its own product" objection.

3 · Objective transcript metrics

4 · Self-reported confidence (in-app, pre + post)

We gauge the students' own confidence at the start and at the end. The question lives inside the Speaking Check — one screen, one tap — so collection is 100%:

Secondary evidence only, but "average confidence went from 4 to 7" is a line investors remember.

5 · What investors see

One page per cohort

  1. Headline: X of Y students improved their blind-rated speaking score; average change +Z/15 (per-criterion too — expect fluency to move most in 3 weeks; say so upfront).
  2. Objective line: average WPM week 1 vs week 3 from in-app attempts.
  3. Confidence line: self-rated confidence pre vs post.
  4. Buy-in line: % of parents who chose to pay RM15/month to continue after the free trial.
  5. The clincher: 3–5 paired audio clips — same student, same task, day 1 vs day 21. Requires the parent-letter consent checkbox.
  6. Honest caveats stated — no control group, 3 weeks, small n, self-administered — which is what makes the rest believable. Directional evidence the 500-student trial confirms.
Structural point for investors: because the Speaking Check is in the product, this exact measurement runs at Chung Hua's 500-student scale with zero field operations — and post-launch it doubles as the conversion mechanic ("hear your day-1 self vs. today") for every trial user.

Logistics & cost

The build

The one engineering deliverable — ships before kickoff, ~Aug 16

A new in-app activity, deliberately thin — it reuses the existing recorder → R2 upload → transcription → attempts machinery:

Checklist