The shape of it
Five functions, one direction of travel, and the two that gate the far end.
02The build
Eighteen systems, mapped onto one student’s path through a session.
03The research program
16 aims, 21 registered hypotheses, six executed studies and what each found.
04The contract
How research and build negotiate — four rounds, and who won each.
05What we own
Eight corpora, their licences, and three measurements that reverse the brief.
06The critical path
Two tracks, five blockers, one convergence. Click a blocker to see what it releases.
07Open decisions
Fifteen calls reserved for a person. Several are load-bearing.
trackable 08Outside parties
Seventeen conversations owed. All drafted, none sent.
trackable 09The cleanup queue
Twelve places two documents disagree. Two would corrupt a study.
10The outside view
What is real, what is still a bet, and what would make us stop.
What the research has established
No located study establishes that prerequisite sequencing works for writing, in either direction. The null is the operative default and the burden is on us — which is why cohort 1 is framed as an estimation cohort rather than a confirmation.
Writing development is asymmetric-reciprocal, not strictly hierarchical, and the componential effect fades with age: transcription→quality runs β = .60 at grades 4–6 but .26, n.s., at grades 7–9. Grade 9 is the hard case for this thesis.
Two trained humans agree at about 0.5 on analytic traits, not 0.8. Across 23 raters: Overall 0.511, Vocabulary 0.394. The founding brief's per-trait bar is not a bar anyone has cleared, humans included.
Under holistic scoring, word count alone predicts the human score at ρ = 0.733. Under trait rubrics it falls to 0.26. Trait scoring is not a finer picture of sub-abilities — it is what stops the scorer rewarding length.
The deployed scorer resists padding — −0.31 per trait, where the statistical baseline gains +0.61. It does not resist fluent nonsense, which still scores 2.60 against 1.36 for genuinely weak writing. The through-line check catches 4 of 6.
Under the rule as originally built, 25% of a struggling student's node encounters were marked ready while true success probability was below 0.70 — the gate was firing by chance. That produced unlock rule v1, at the cost of roughly an 80% slowdown.
No node is too big or too narrow to teach. The session is too short. Share of struggling students finishing the module: 22% at 30 minutes, 60% at 35, 93% at 40.
Across 4,992 real writers, process data predicts quality only through fluency and length. It is authorship evidence, not a scoring input. The existing 40% paste rule sits above the 99th percentile of honest behaviour, and the flags do not fall harder on ELL writers.
In the closest public corpus of AI-authored items, 18.7% had a broken answer key — and that is a floor. Two reviewers is the right number: new defects found run 1.33, then 0.97, then 0.53.
Automated presence detection cannot grant credit — counterclaim recall is 0.036 at 85% precision. Run backwards it earns its place: when it says an element is absent it is right 93–96% of the time.
The shape of it
Five different kinds of work run inside one project, and four of the five sit upstream of a pilot that has not started. This is the order of operations.
Engineering shop
Student app, engine, scorer, backend, grading tool. Ships fast — the app is eighteen bundle versions past its first release. Deployed on Squarespace with a Supabase backend shared with the VPA dashboard.
Research lab
A nine-stage agent orchestration (~1,300 lines, 25 tests) running literature streams, hypothesis generation, red teams and experiment design. Its output is a frozen registry and an approved experiment.
Content studio
965 tasks across 18 passages, AI-authored at ~63 tasks per passage in ~10 minutes of agent time. Zero human-reviewed. Expect roughly one item in five to have a broken key.
Compliance function
FERPA, COPPA, Colorado's student-data act, consent records enforced in row-level security, retention enforced in code. The software side is unusually far ahead; the paperwork side has not started.
Capital & channel
SGO analysis, family-directed funding, D2C-vs-district channel, rater recruiting, investor relations. Runs against external dates you do not control — the federal tax credit goes live 1 Jan 2027.
Nobody owns the seam
Fourteen documented contradictions between documents — three different ASAP baselines, seven vs eight traits, an experiment still written for grades 6–8. No one is assigned to reconcile them. See §09.
The build
Eighteen distinct systems exist. The clearest way to see them is as the path one student takes through a single 35-minute session — and the four places where that path runs into something that has not been built.
The eighteen systems
| System | Version | State | The number that matters | Biggest open risk |
|---|---|---|---|---|
| Student app — Elevate Writing | v1.2 / -20.js | Deployed | 965 tasks render, 0 key failures | Manual Squarespace deploy; no document states which bundle is live right now |
| Task bank | v0.4.2 | In production | 965 tasks / 18 passages | Zero human-reviewed. Expect ~19% to have a broken key |
| Backend — Supabase | v0.6.0, schema v2 | Stub-tested only | 9 tables, RLS enforced | Never exercised end to end with real accounts |
| Scorer — score-writing edge fn | Sonnet 4.5 + rubric | Live | 40 items, 0 errors, ~18s/call | Its agreement with human scores has never been measured. Bar is 0.786, not 0.668 |
| Through-line check | deployed | Mis-configured | threshold 0.50, should be 0.70 | False-flag rate on real student writing is unknown — measured on 6 synthetic items |
| Unlock / mastery engine | rule v1, frozen | Built, 11 tests pass | 13.7 attempts per node (was 7.6) | ~80% pace slowdown. Struggling tercile needs all 18 sessions. Never tested on a human |
| Instruction model | v0.1 | Browser-verified | 4 strategies × 4 fading stages | Pure hypothesis. The top rung's concession step has no worked example authored |
| Teacher grading tool | v0.2 / v0.5 review flow | Deployed, no data | 7 traits, blind then reveal | Zero human ratings exist. This tool is where the calibration set would land |
| Spaced review scheduler | randomised bands | Built | half-lives 1.3–1.9 days at every level | Per-skill intervals are unidentifiable from a one-month module. Probes need 30/60/90 days |
| Pilot pathway — the graph the engine runs | v0.3, frozen | Approved & frozen | 13 nodes / 14 edges / 5 levels | All 14 edges still typed uncertain. Two-teacher edge review never returned |
| Long-range skill graph | v0.2 deployed / 188 designed | Stale | 203 live vs 188 designed vs 13 running | No version on the site matches the design, and nothing maps the 13 pilot nodes onto the 188 |
| Offline baselines — ASAP + W-Pal | v0.1 / v0.2 | Trained | QWK 0.786 → 0.465 within length decile | Everything is length. tf-idf gains +0.61 when you paste the essay's own sentences back in |
| Presence detector — PERSUADE | v0.1 | Role redefined | counterclaim recall 0.036 at 85% precision | Cannot grant encompassing credit. Useful only backwards, as an absence screen |
| Adversarial suite | 80 offline / 40 live | Built and run twice | nonsense scores 2.60 vs 1.36 for weak writing | Fluent nonsense is still the hole. Through-line at 0.70 catches 4 of 6 |
| Presence rubric + rater set | v2, R1–R67 | Agent-tested only | agent–agent κ 0.95–1.00 | That κ satisfies nothing. 8 gold labels need human adjudication before the queue can open |
| Shadow-trial dashboard | 24 seeded responses | Live, unrun | needs 5–10 teachers, then 50–100 | No teacher-appropriateness data exists |
| Session event stream & audit log | — | Not built | 7 timing field groups, all new | None of the packet's timing requirements is derivable today |
| Experimental-arm machinery | — | Not built | REQ-01, REQ-02, REQ-06 | The primary confirmatory test of the whole program currently has no engine. |
The research program
Three numbering systems run at once. They look like alternatives and are in fact layers: A2-3 is a coordinate — aim A2, bullet 3 — not a separate experiment. Read as layers, the program is compact.
Program aims
Five themes (Ontology, Measurement, Instruction, Adaptation, Implementation) holding 16 aims, each with a numbered to-do list. “A2-3” means aim A2, bullet 3 — not a separate study.
Hypothesis registry
21 hypotheses in six families, frozen at v3. Each names a prediction, a falsifier, and a sample size. One primary confirmatory test; no single result may edit the graph.
Runs
Six executed pieces of work, each with a ledger, verification counts and an approval record. This is the layer with findings in it.
What actually ran, and what it found
| Run | Question | Headline finding | Did it change the build? |
|---|---|---|---|
| Cycle 1 37 agent runs, 9 stages | Does diagnostic-driven prerequisite sequencing beat a fixed sequence? | No direct evidence exists either way in any located study. 119 claims from 77 publishers; 106 sourced, 81 independently verified, 9 disputed and corrected. H0 is the operative default; the burden is on H1. | Yes — reframed cohort 1 from a confirmation study to a preregistered estimation cohort |
| A1-8 reciprocity stream | Is writing development hierarchical or reciprocal? | Asymmetric-reciprocal, and it fades by grade 9. Transcription→quality β=.60 at grades 4–6 but .26 n.s. at 7–9. Grammar drill ES −0.32 vs strategy 0.82. 12 new claims, 12/12 verified. | Yes — produced pathway v0.3: dropped CLAIM-01→PROC-01 and SENT-01→SENT-02, moved PROC-01 to level 4 |
| A2-4 session-time audit | Is any node too narrow or too big for a 35-minute session? | No node is wrong — the session is too short. All-12-ready by session 18: at 30 min 98/79/22%; at 35 min 100/99/60%; at 40 min 100/100/93%. Unplanned finding that mattered more: 12–20% of struggling students were being marked ready by chance. | Yes, though indirectly — the false-ready finding is what produced unlock rule v1 |
| B1-2 presence rubric | Can presence-with-meaning-preserved be coded reliably? | Rubric v2 (67 rules) reached agent–agent κ 0.95–1.00 on five of six codes. The run states plainly that this is an upper bound on rubric clarity and satisfies nothing in MEAS-01. 12 disagreements on 8 items need human adjudication. | Yes — three packet requirements, including the insertion flag as a database-generated column |
| A2-3 prompt stability | Is a placement a property of the student or of the passage? | Designed, red-teamed, item-authored, readability-corrected, independently validated. Honest power: at n=70, INDETERMINATE is the modal outcome for true agreement 0.73–0.87. ~70–80 adult hours. | No — and nobody is waiting for it. Explicitly on the do-not-build list, aimed at a registry v4 that does not exist |
| KLiCKE baseline | What does honest human composition look like keystroke by keystroke? | 4,992 writers. Median 345 words in 29 min; 57% of the session producing nothing; pasting median 0. Process predicts quality only through fluency and length — so process is authorship evidence, not a scoring input. Your existing 40% paste rule sits above the 99th percentile of honest behaviour, and the flags do not fall harder on ELL writers (5.0% vs 5.7%). | Yes — three shipped product changes |
| WAT instrument | Should a writing-apprehension measure go in? | Daly-Miller 26-item adopted, verbatim from ERIC ED345265, ~5 min, before both cold prompts. KTEA-3 put on hold — off-construct for transfer, ~$1,147 for two forms, 30 proctored minutes each. | Yes — decision dec_d4618a8448, packet REQ-20 |
The one test that decides everything
SHAPE-03 is the primary confirmatory hypothesis, and inside it sits the product thesis: does module-end mastery predict delayed unaided writing?
- Needs ~190 analysed per arm. Cohort 1 at 30–80 students gives a minimum detectable effect of d ≈ 0.61–0.99 — an estimate, not a verdict, and the approval record says so.
- The success metric was rebuilt after simulation showed the original had 12% power: counting mastered nodes correlates with delayed writing at r = 0.12 and needs N≈535. The replacement — mean success rate over the module's final third — correlates at r = 0.49, giving 80% power at N = 31, and correctly drops to 0.15 in a simulated “Quill world”.
- The kill criterion is written down. If partial r < 0.10 with a CI upper bound below 0.20 across two cohorts, the componential thesis fails for this grade band and the product pivots to composition-first.
Work running with no consumer
Worth knowing before you fund more of it. None of this is wasted — but none of it is currently blocking or unblocking anything.
- The prompt-stability study (A2-3). Fully designed and validated, then placed on the do-not-build list, aimed at a registry version that does not exist.
- Edge annotations v2 and v3. All 14 edges still typed uncertain; the brief states human review of them blocks nothing in cohort 1.
- The scorer research. The 80-item adversarial set, the bias tables and the evaluation harness have no live input while the calibration set does not exist.
- The second research cycle. Seven stream briefs were scoped; only the reciprocity rerun was executed. Stream 1 — time-equated sequencing — is the one that would actually move the H1 evidence base, and it has not run.
The contract between research and build
The most unusual piece of the program, negotiated across four rounds. Research does not send prose to build. It emits a machine-assembled implementation packet; a building agent must validate it and refuse an unapproved one, then return a traceability report mapping every requirement to where it was implemented and tested.
| Round | Who moved | What happened |
|---|---|---|
| 0 | Research → Build | Cohort 1 does not depend on live model scoring — the primary outcome is adjudicated human double-rating; the scorer is shadow-only. Composition is a fixed 12-minute slot with an 8-minute floor and a hard cap, value to be decided by a human. |
| 1 | Build → Research four contradictions | Answered from the code, and corrected research on things that changed the arithmetic: the ready rule was 2-of-3 with retries counting, not 3-of-4 — so the false-ready figure was a floor, not an estimate. The engine computed a transfer-anchored verified state and gated nothing on it, which research called “the Quill failure mode written into the unlock logic”. Composition was already rung-based (3/6/9/14), not 12 minutes. And no arm concept existed at all. Build's closing line: “Nothing here is a decision. It arrived as a document from an agent run, not as a decision from Samuel Bennett.” |
| 1b | Build → Research the blocker | The two graphs were not the same object despite both having 13 nodes and 16 edges. The quotation dependency ran in opposite directions — the exact pair two registered hypotheses were designed to test. Verdict: nothing should be frozen until one is designated authoritative. |
| 2 | Research → Build concede three | Accepted all three corrections and re-ran the simulation under the real rule: 25% of node encounters marked ready while true success probability was below 0.70 (36% on SENT-02). Conclusion: “the rule is not slower, it is just wrong more often for weak students.” Three recommendations routed to you. |
| 3 | Build decides one, you decide two | Build refuted research's unlock recommendation on arithmetic: a disjunction can only unlock more often than its strictest branch, so “provisional OR first-try CR” scored 20.6% false unlocks — worse than the rule it replaced. Unlock rule v1 frozen instead at 12.9%. Your two decisions: dec_f353a96618 (variable slot, 30-XP goal, no cap at all) and dec_09f45341d4 (pathway v0.3 is the graph). |
| 4 | Research re-derives | Yoking moved from quantity to rule. Both arms must now run a byte-identical config, hashed per session and audited daily; composition minutes are expected to differ by arm and are reported rather than policed. The causal claim narrows from “sequencing at equal dose” to “the policy as deployed”. awaiting your sign-off |
Superseded
Graph 0.2-reconstructed, 16 edges. 25 requirements, 24 telemetry rows. Wrong graph, wrong dose model, wrong mastery rule.
Superseded almost at once
The substantive rewrite: pathway v0.3, yoking by rule, the 30-XP session goal, unlock rule v1, and the CR bank gap as a P0 requirement. 27 requirements.
The live contract
28 requirements (19 are P0), 32 telemetry rows, 16 guardrails, 14 prohibited changes, 9 open human decisions, 14 items on do-not-build-yet. Only change from v5: the WAT.
What the program actually owns
Separate from the code: a set of licensed corpora, a set of measurements nobody else has run on them, and a task bank. The measurements are the underrated asset — several of them reverse things the founding brief asserts.
| Corpus | Size | Licence | What you may do with it | State |
|---|---|---|---|---|
| ASAP 2.0 | 24,728 essays grades 6–10 | CC BY 4.0 | The only commercially trainable corpus. Carries ELL, economic, disability, race and grade labels on every essay — so it also carries all bias testing | Trained · QWK 0.786 |
| PERSUADE 2.0 | 84,440 spans 8,426 essays | CC BY 4.0 (Lab release) | Argument-element detection. Licence reversal: the GitHub copy is BY-NC-SA; the Learning Agency's own release is CC BY — and that is the copy you hold. The brief's “pull from GitHub not Kaggle” rule is backwards here | Ships as absence screen |
| KLiCKE | 4,992 writers | CC BY 4.0 | The honest-human typing baseline behind every authorship flag | Baseline computed |
| ELLIPSE | 8,890 essays 23 raters | CC BY-NC-SA | Nothing trained on it may ship. ShareAlike is the clause that could reach into the product, not just the input. Measurement only | Measured, cannot ship |
| College Readiness Math | 434 review notes 321 items | CC BY-SA or CC BY (conflict) | Mined into your 19-defect review taxonomy and the reviewer marginal-return curve | Mined |
| PIILO | 6,807 docs | CC BY 4.0 | Would evaluate the PII scrubber upstream of the model | Not started |
| AIDE / Quest | 1,378 / 6,989 | CC BY 4.0 | AI-detection check; reading-comprehension taxonomy | Deprioritised / shelved |
Three measurements that change the brief
- Two trained humans agree at about 0.5 on analytic traits, not 0.8. Across ELLIPSE's 23 raters: Overall 0.511, Grammar 0.463, Vocabulary 0.394. The founding brief's QWK ≥ 0.80 per-trait bar is not a bar anyone has cleared, including humans. The replacement: holistic keeps 0.80 against a two-rater average; per trait the bar becomes model–human agreement at or above human–human on the same essays.
- The six traits are one thing measured six times. First eigenvalue 73.3% of variance; second 0.377. Expect organisation, sentence clarity and conventions to collapse into one language-control trait. The four argument traits are a different construct — and nothing public validates them, which is precisely why your own students are the only route.
- Analytic rubrics defeat the length effect — that is the real case for traits. In ASAP, word count alone predicts the human score at ρ = 0.733. In ELLIPSE, trait scores run 0.093–0.304. Not a finer picture of sub-abilities; a scorer that stops rewarding length. That is the property to protect and the argument to make.
And one about the task bank
From 434 expert review notes on 321 AI-generated items — the closest public analogue to your own authoring pipeline:
- 18.7% had a broken answer key or a duplicate correct answer, and that is a floor. “An AI-authored task bank is not a draft that needs polishing, it is a draft where roughly a fifth of the items are wrong on the thing that matters most.”
- Two reviewers, not one and not three. New defects found by reviewer 1/2/3/4: 1.33, 0.97, 0.53, 0.25.
- Reviewers barely agree on what is wrong. Mean overlap of defect sets 0.26; on “is the key broken”, 35% agreement. Budget human attention for skill alignment and model-answer correctness; automate rendering, house style, reading level and answer-position balance.
Applied to your 965: about 180 items with a broken key, and ~1,468 review passes to find them.
The critical path
Five things block the pilot. They are not equally urgent, because they are not all on the same track — and the one people reach for last is the one that is strictly serial and takes the longest.
The five blockers — click one to see what it releases
Decisions awaiting a named human
Agents wrote almost everything in this program. These are the items the design explicitly reserves for a person to decide, plus the commercial calls that go with them. Several have been open since the program began and are now load-bearing. Click a status to cycle it — the marks are shared with everyone who opens the page.
Product and research
Named by the experiment designer as a blocker — every XP-derived threshold is parameterised on it. Under current values a rung-1 session reaches the 30-XP goal in 4–6 active minutes and a rung-4 session in 16–20. If students stop at goal, 18 sessions yield as little as 90 instructional minutes with a 53% assessment share. The 45/30/25 task mix cannot fit inside the goal.
The v7 re-derivation replaces equal-dose yoking with byte-identical config hashed per session. It also narrows what you can claim from “sequencing at equal dose” to “the policy as deployed”. Everything downstream — what FEAS-01 means, the fidelity thresholds, the preregistration hash — waits on this.
Open from the start, and more load-bearing now that hours differ by arm. The registry says the ~28-minute diagnostic and verification minutes count with a sensitivity exclusion; the experiment's analysis plan excludes them. Idle threshold is 60 s in one document and 90 s in the other. Both are live.
S07, S16, S20, S21, S26, S30, S32, S36 — twelve disagreement records where both raters went against gold in the same direction. Held out of the rater-training queue until a named human rules. This blocks the training set, which blocks rater training, which blocks MEAS-01, which blocks everything.
The founding brief still states it; ELLIPSE shows two trained humans reach ~0.5 on analytic traits; the W-Pal work declares it unreachable by these methods; the backend doc still cites it as the operative rule. Four documents, three positions. Proposed replacement: holistic keeps 0.80 against a two-rater average, per trait the bar becomes model–human ≥ human–human on the same essays, plus stability under attack.
“Fix the primary outcome and analysis plan in writing before any student is enrolled” appears as a step in the explainer and as requirement S1 in the research program. It has no signature. Must include retention probes at 30, 60 and 90 days past the module — a one-month module's randomised intervals produce a median realised delay of about a day and a half.
plan_before_draft and goal_named need new UI, which changes the frozen composition structure — so it must be decided before the freeze, not after. Research's own position is that revision-event count and minutes-before-first-draft-text suffice for cohort 1.
An edge claim is three separable claims — dependency validity, instructional efficiency, diagnostic predictiveness — reported separately. No format exists, so EDGE results have nowhere to land in the change log. Related: the two-teacher edge review form is still outstanding, and its six flagged edges are supposed to replace the six the registry is currently frozen around.
Students in the diagnostic arm finish earlier and get extension work. Either cap that exposure, or disclose that H1 is being tested as a package rather than as sequencing alone. None exist in the engine either way.
The answer sheet prints the student's code in a box and a photograph can contain a name the student wrote there. The PII scrubber cannot see inside an image. Either transcription becomes the redaction point, or handwriting comes out of the model path until it can be checked.
Commercial and legal
Closed: Team America Foundation cannot be an SGO as structured — it is a private operating foundation, and “operating” is a subtype, not an exemption. Converting it is explicitly advised against. Open: Path A (register VPA as an eligible provider — free, do it regardless), Path B (partner with ACE Scholarships — one phone call, but they set the award criteria), Path C (a Vail Valley community scholarship fund — only after Treasury finalises the disqualified-persons rules).
Nonprofit status, bridge financing from private partners, a platform with investor-style economics, and a related scholarship organisation potentially funding your own students. These four interact, and the recommendation is one session with a nonprofit tax attorney covering all of them at once rather than four separate conversations.
Crossley's lab may well reframe a paid-hours request as a research collaboration with co-authorship. Say “paid contract hours” explicitly in the first email if that is what you want. Your target hires are exactly the people who would want to write a paper about this — which is also why the contractor agreement needs an explicit IP assignment and a no-external-publication clause.
Three benches, and conflating them is the failure mode. $30–40/hr for raters is genuinely generous — Measurement Incorporated pays $15, ETS about $25. But $30–50 is below market for the expert tier: budget $50–75/hr or flat fees of $1,500–5,000. The outreach templates still carry bracketed placeholders.
$500–2,000 per student per semester fits inside ESA award sizes and inside a $1,700-granularity scholarship pool, and needs no purchase order, RFP or procurement cycle. Open question worth money: whether an accredited, credit-bearing course prices into a better-funded ESA bucket than “tutoring”.
Outside parties, and what each is owed
Every one of these is drafted and none is sent. The outreach templates are written, the addresses are verified against live pages, the sequence is planned. The gap between this program and its next phase is largely a set of emails.
FERPA's school-official exception runs through the school, not through you — you cannot self-designate contractors as school officials. De-identification handles rating; the formal studies exception is what you need to link a student's baseline essay to their delayed-transfer essay, which the success metric requires. Start it now; it will take longer than recruiting. Contacts: Cat, Jared, Meg.
Names the school-official designation, the authorised purpose, the ban on redisclosure and on training anyone's models, the retention periods and the deletion process. Until it exists, no student account should be cleared for consent. Alongside it: the published privacy notice carrying retention numbers verbatim, and the written parental-consent process.
He built ASAP 2.0, PERSUADE 2.0 and ELLIPSE — the three corpora your whole scoring plan rests on. The lab's jobs page explicitly offers paid student positions with no formal application; the lab is rebuilding after a move from Georgia State, so there is less competition for student hours. It also publishes on bias in LLM writing feedback. scott.crossley@vanderbilt.edu
Best Colorado rater pipeline: she directs English Education and coordinates the MA, whose students are current and aspiring ninth-grade English teachers — the exact rater profile. Two hours from Edwards for in-person calibration. stacy.bailey@unco.edu
Co-created the Intelligent Essay Assessor and ran large-scale automated scoring at Pearson. The most credentialed automated-essay-scoring person in Colorado, and has the industry instincts to understand a paid consulting ask without a grant wrapper. Aim him at adversarial robustness. peter.foltz@colorado.edu
A whole doctoral programme trained in rater agreement, validity, scaling and subgroup analysis; one email to the director reaches the cohort. Also there: William Penuel on research-practice partnerships. derek.briggs@colorado.edu
$3.5M+ in Gates, NSF and Spencer funding on automated writing evaluation efficacy. His research question is your governing risk. Small PI-centred group, fast personal replies. joshwils@udel.edu
She built Writing Pal, the closest historical analogue to this product including its own transfer lessons. You have already rebuilt 28 of its indices and measured where they fail. Also at ASU: Maria Goldshtein, an “Inclusive Language Analytics” postdoc — a near-exact match for the subgroup-fairness reporting requirement.
Four things to get: the economics of a fine-tuned classifier versus LLM calls at your volume and latency (your open scorer-architecture question, and he has worked it for a consumer product at scale); whether Barak Oshri or Violet X. would review the trait-scoring design — you have no ML advisor and QWK-with-subgroup-reporting is a real ML problem; what he learned from 100+ interviews about why people don't write; and the channel conversation. Be explicit that his success metric is engagement and yours can never be — an investor will conflate the two.
Denver-based, already operating in Colorado, explicitly markets SGO infrastructure. Foundation supporters donate and take the credit; VPA families apply. No entity to form, no 90% rule, no public support test, no annual state certification. The trade is that ACE decides who gets funded. Must happen in 2026, before the 1 January 2027 start.
Team America Foundation cannot become an SGO as structured, and converting it is an expensive way to get a worse version of the ACE partnership. This needs to go back to the foundation as a finding, not sit in a document. Also confirm whether TAF files Form 990-PF — that is the definitive confirmation of private-foundation status. TAF could donate to a future SGO, but a private foundation granting to an organisation it effectively controls raises self-dealing questions.
Nonprofit tax: entity structure, bridge financing, investor-style platform economics and a related scholarship org, covered together. Privacy and education: review the schema and privacy document, rule on whether Colorado's student-data act adds anything beyond FERPA and COPPA, review the contractor agreement template once so it can be reused, and confirm the corpus licences before anything ships.
No training on commercial API data by default, but zero-retention eligibility is unconfirmed. Get it in writing and attach it to the DPA as the subprocessor record, before the pilot.
Still outstanding. Six edges are flagged suspect and undecided, and the registry's six EDGE hypotheses were chosen by graph topology rather than by anyone who teaches ninth grade. The form exists and is ready to hand over.
The tool is live with 24 seeded responses: blind rubric scoring, then reveal, then helps / should-defer / show-feedback verdicts, with teacher–model and teacher–teacher agreement computed automatically. Expands to 50–100. Produces the only teacher-appropriateness data that will exist.
FERPA-aware confidentiality, explicit IP assignment, no-external-publication, 1099 classification safeguards. Counsel reviews it once; reuse forever. Watch three things that will bite: F-1 candidates need written DSO confirmation before signing, funded RAs face outside-work caps, and W-9 before first payment.
Steward of ASAP 2.0 and PERSUADE going forward, and the source of the licence discrepancy you found. Worth a conversation regardless. Perpetual Baffour there works on bias in automated essay scoring. info@the-learning-agency.com
The cleanup queue
Work at this pace leaves a wake. These are places where two project documents disagree, or where a file says something that is no longer true. None is fatal; several would not survive an accreditor’s read, and two would corrupt a study if left alone.
| What disagrees | The two versions | Why it matters |
|---|---|---|
| The rater training set contradicts itself | Nine items carry v2 gold labels with v1 rationales still attached, and S30's rationale cites a rule that v2 retired | Fix first This file seeds the rater training queue. Trainees would read reasoning that contradicts the labels they are scored against |
| Two Gettysburg transcriptions | The rater set ships the Gutenberg text (“cannot”, “from this earth”); the rubric anchors quote the Bliss forms (“we can not”, “from the earth”) | Fix first A rater checking a quotation against the passage copy will hit a false mismatch on a rubric about quotation accuracy |
| Three different ASAP headline numbers | QWK 0.807 (n=7,421 official split) vs 0.786 (80/20 by essay). Length baselines 0.724 / 0.668 / 0.652 | The readiness gate is stated as “beat 0.786”. Which split is canonical is not settled anywhere, and the difference decides whether the scorer passes |
| Subgroup bias points two directions | One document reports signed bias of −0.110 for ELL and −0.149 for disability and concludes lower-scoring groups get pushed further down. Another reports mean signed error within 0.05 of zero for every subgroup and concludes the gaps are in agreement, not in level | Different models on different splits — but the documents do not say so, and only one of the two can be quoted to an accreditor |
| Seven traits or eight | The rubric spec says eight traits with 0–4 anchors; the grading tool scores seven; the training doc says seven | Nothing reconciles it. The missing one appears to be independence, which is human-only — but that should be written down, not inferred |
| The experiment is still written for grades 6–8 | Grade 9 is decided; the experiment's population, moderators, strata and inclusion criteria still read 6–8, and the MDE figures still need correcting to 0.61 / 0.77 / 0.99 | The preregistration hash cannot be taken over inconsistent text. Also pending: one claim needs relabelling as unverified, and an instrumentation field still describes the retired fixed slot |
| Pathway size is stated four ways | 188 nodes (brief) · 203 (deployed at /graph) · 34 (engine v0.1) · 13 (what actually runs) | Nothing maps the 13 pilot nodes onto the 188, and the brief calls for a 25–35-node induced subgraph. The 34-node version seems abandoned without being declared so |
| Registry text describes a retired design | FEAS-01 still says “8-minute floor / hard cap”; standing rule 4 still says composition minutes are yoked; the cohort scope still says “feasibility of equalizing writing time” | All three were superseded by a later decision. They need a human-approved amendment, not a quiet edit |
| Approved files still say “proposed” | Every change record inside the approved pathway file still reads proposed (designed, not decided) though the file-level status is approved and frozen | Minor, but it is the provenance trail an accreditor reads |
| The site map is eighteen bundles stale | It points /writing-app at bundle -2; the current one is -20. No document states which bundle is live on Squarespace right now | It is also the only document describing the review page, the 203-node graph and the legacy diagnostic |
| Task review counts disagree | 965 pending in one document, 734 in another; 232 legacy tasks versus an implied 231 | Small, but it is the denominator of your largest content job |
| Corpus sizes disagree with the brief | ASAP 24,278 (brief) vs 24,728 (held); ELLIPSE “~6,500” (brief) vs 8,890 collected / 6,468 published — and the brief groups it with the commercially usable sets when it cannot ship | The ELLIPSE one is a licensing error in the founding document, not a typo |
The outside view
Everything above is the operator’s view. This is the version for someone coming to it cold: what is real, what is still a bet, what the next round of resourcing buys, and what would make us stop.
A working platform and a research organisation
- A deployed student app running a 13-skill pathway over 965 authored tasks, with placement, adaptive selection, spaced review, a composition every session, and a decision log that records a written reason for every choice it makes.
- A scoring function live in production that resists the attacks that beat the statistical baseline — padding an essay by 40% costs it a third of a point per trait, where the conventional model gains 0.61.
- A backend with consent enforced in the database rather than the interface, retention enforced by code, and a teacher grading tool that computes agreement automatically.
- A research program with 21 registered hypotheses, a frozen registry, a written kill criterion, and a machine-readable contract that a building agent must validate before it implements anything.
All of it, honestly — and that is the design
- No student has used this. Every number on this page comes from public corpora, simulations or agent runs. The simulations are good — they caught that the original design produced zero compositions in twenty sessions, and that the original success metric had 12% power — but simulated students are not evidence about real students, and the documents say so on every page.
- Nobody has shown that prerequisite sequencing works for writing. The literature cycle's conclusion was that no direct evidence exists in either direction, so the default is that it doesn't, and the burden is on us. That is an unusually honest place to start and it is why cohort 1 is framed as an estimation cohort rather than a confirmation.
- The scorer has never been checked against a human. Not because it failed — because the measurement has not been run.
What the next round of resourcing buys
- A rater bench. Two trained humans on 100–150 real ninth-grade paragraphs at $30–40/hr, plus one or two expert consultants at $50–75/hr. This single item unlocks every measurable claim in the program, and it is the difference between “the model agrees with itself” and “the model agrees with teachers”.
- Content review hours. ~1,468 review passes over 965 tasks, where the public evidence says one item in five has a broken key.
- ~280 new constructed-response items so that five of the thirteen skills can be gated on writing rather than on selection.
- Engineering for the experimental arms — currently the primary test of the whole thesis has no engine.
- Two single-session legal engagements, one tax and one privacy.
What would make us stop
This is written down and it is unusual to have:
The whole design exists to get that answer in month three rather than year three. The documented failure of this product category is transfer — drill gains that never reach real writing. Quill's own randomised trial moved a sentence-combining test by d = 0.30 and moved actual composition by nothing distinguishable from zero. Every structural decision here — composition from session one, mastery defined as predicted use in free writing, more writing rather than more drill when the two diverge — exists to avoid rebuilding that product.