Writing Academy · Operating Map

The whole program,
laid out end to end

Working reference
assembled from 52
project documents

This program runs five organisations at once — an engineering shop, a research lab, a content studio, a compliance function and a capital raise — and they hand work to each other in a fixed order. This is that order, with the current state of every piece. Everything below is drawn from the project’s own documents; where two of them disagree, both numbers are shown.

18systems built or deployed
965tasks authored
21registered hypotheses
28build requirements in the live packet
5blockers gating the pilot
0human ratings collected
Contents — jump to any part
Findings so far

What the research has established

Full research program → 10 findings
Sequencing

No located study establishes that prerequisite sequencing works for writing, in either direction. The null is the operative default and the burden is on us — which is why cohort 1 is framed as an estimation cohort rather than a confirmation.

Development

Writing development is asymmetric-reciprocal, not strictly hierarchical, and the componential effect fades with age: transcription→quality runs β = .60 at grades 4–6 but .26, n.s., at grades 7–9. Grade 9 is the hard case for this thesis.

Measurement

Two trained humans agree at about 0.5 on analytic traits, not 0.8. Across 23 raters: Overall 0.511, Vocabulary 0.394. The founding brief's per-trait bar is not a bar anyone has cleared, humans included.

Why traits at all

Under holistic scoring, word count alone predicts the human score at ρ = 0.733. Under trait rubrics it falls to 0.26. Trait scoring is not a finer picture of sub-abilities — it is what stops the scorer rewarding length.

Scorer

The deployed scorer resists padding — −0.31 per trait, where the statistical baseline gains +0.61. It does not resist fluent nonsense, which still scores 2.60 against 1.36 for genuinely weak writing. The through-line check catches 4 of 6.

Mastery gate

Under the rule as originally built, 25% of a struggling student's node encounters were marked ready while true success probability was below 0.70 — the gate was firing by chance. That produced unlock rule v1, at the cost of roughly an 80% slowdown.

Time budget

No node is too big or too narrow to teach. The session is too short. Share of struggling students finishing the module: 22% at 30 minutes, 60% at 35, 93% at 40.

Authorship

Across 4,992 real writers, process data predicts quality only through fluency and length. It is authorship evidence, not a scoring input. The existing 40% paste rule sits above the 99th percentile of honest behaviour, and the flags do not fall harder on ELL writers.

Content quality

In the closest public corpus of AI-authored items, 18.7% had a broken answer key — and that is a floor. Two reviewers is the right number: new defects found run 1.33, then 0.97, then 0.53.

Encompassing credit

Automated presence detection cannot grant credit — counterclaim recall is 0.036 at 85% precision. Run backwards it earns its place: when it says an element is absent it is right 93–96% of the time.

01

The shape of it

Five different kinds of work run inside one project, and four of the five sit upstream of a pilot that has not started. This is the order of operations.

Evidence base 8 corpora, 131 verified literature claims Research program 16 aims · 21 hypotheses 6 executed studies Build app, engine, scorer, bank, backend Cohort 1 pilot grade 9 · 30–80 students not started reads packet v6 28 reqs serves evidence graph → registry → graph change log Consent & compliance DPA · privacy notice · parental consent none executed gates enrolment Capital & channel SGO path · ESA rails · accreditation runs on its own clock → Jan 2027
Five functions, one direction of travel. The accent arrow is the unusual part: research hands build a machine-readable contract, not prose. The dashed return loop is the only thing that turns the pilot into a research asset — and the red gate is what stops the pilot starting at all.
Function 1

Engineering shop

Student app, engine, scorer, backend, grading tool. Ships fast — the app is eighteen bundle versions past its first release. Deployed on Squarespace with a Supabase backend shared with the VPA dashboard.

Function 2

Research lab

A nine-stage agent orchestration (~1,300 lines, 25 tests) running literature streams, hypothesis generation, red teams and experiment design. Its output is a frozen registry and an approved experiment.

Function 3

Content studio

965 tasks across 18 passages, AI-authored at ~63 tasks per passage in ~10 minutes of agent time. Zero human-reviewed. Expect roughly one item in five to have a broken key.

Function 4

Compliance function

FERPA, COPPA, Colorado's student-data act, consent records enforced in row-level security, retention enforced in code. The software side is unusually far ahead; the paperwork side has not started.

Function 5

Capital & channel

SGO analysis, family-directed funding, D2C-vs-district channel, rater recruiting, investor relations. Runs against external dates you do not control — the federal tax credit goes live 1 Jan 2027.

Missing function

Nobody owns the seam

Fourteen documented contradictions between documents — three different ASAP baselines, seven vs eight traits, an experiment still written for grades 6–8. No one is assigned to reconcile them. See §09.

What it adds up to. There is now a complete adaptive writing platform, a research organisation built to test it, and the legal and commercial case for both — and every one of them is waiting on the same two things: two trained humans scoring a hundred real ninth-grade paragraphs, and a signed data agreement.
02

The build

Eighteen distinct systems exist. The clearest way to see them is as the path one student takes through a single 35-minute session — and the four places where that path runs into something that has not been built.

LIVE PATH DOES NOT EXIST Placement diagnostic 12 items · no early stopping · ~28 min · adaptive form refused at 0.80 vs a 0.90 bar Session planner composition slot reserved first (3/6/9/14 min by rung) · ≤2 due reviews · frontier, ≤3 tasks per skill Attempt one of 965 tasks · 13 task types · 18 passages · level-targeted selection Constraint check auto · most daily volume score-writing traits + through-line share Evidence log writing_submissions · writing_decisions · writing_flags · every decision carries a written reason unlock rule v1 → frontier Arm assignment no table, no flag, no allocator — the primary test has no engine Constructed-response tasks 5 of 13 nodes have none · ~280 items to author Human rating queue 2 trained raters, κ ≥ 0.70 — zero ratings collected Session event stream no timing fields derivable today Everything on the left runs today. Nothing on the right does.
One session, end to end. The loop on the left is the product: the unlock rule decides what opens next, and it now requires four unaided successes and one first-try constructed response — which is exactly why the missing CR tasks on the right matter. Five of thirteen nodes currently fall back to a weaker rule because the bank has nothing to gate them on.

The eighteen systems

Show
SystemVersionStateThe number that mattersBiggest open risk
Student app — Elevate Writingv1.2 / -20.jsDeployed965 tasks render, 0 key failuresManual Squarespace deploy; no document states which bundle is live right now
Task bankv0.4.2In production965 tasks / 18 passagesZero human-reviewed. Expect ~19% to have a broken key
Backend — Supabasev0.6.0, schema v2Stub-tested only9 tables, RLS enforcedNever exercised end to end with real accounts
Scorer — score-writing edge fnSonnet 4.5 + rubricLive40 items, 0 errors, ~18s/callIts agreement with human scores has never been measured. Bar is 0.786, not 0.668
Through-line checkdeployedMis-configuredthreshold 0.50, should be 0.70False-flag rate on real student writing is unknown — measured on 6 synthetic items
Unlock / mastery enginerule v1, frozenBuilt, 11 tests pass13.7 attempts per node (was 7.6)~80% pace slowdown. Struggling tercile needs all 18 sessions. Never tested on a human
Instruction modelv0.1Browser-verified4 strategies × 4 fading stagesPure hypothesis. The top rung's concession step has no worked example authored
Teacher grading toolv0.2 / v0.5 review flowDeployed, no data7 traits, blind then revealZero human ratings exist. This tool is where the calibration set would land
Spaced review schedulerrandomised bandsBuilthalf-lives 1.3–1.9 days at every levelPer-skill intervals are unidentifiable from a one-month module. Probes need 30/60/90 days
Pilot pathway — the graph the engine runsv0.3, frozenApproved & frozen13 nodes / 14 edges / 5 levelsAll 14 edges still typed uncertain. Two-teacher edge review never returned
Long-range skill graphv0.2 deployed / 188 designedStale203 live vs 188 designed vs 13 runningNo version on the site matches the design, and nothing maps the 13 pilot nodes onto the 188
Offline baselines — ASAP + W-Palv0.1 / v0.2TrainedQWK 0.786 → 0.465 within length decileEverything is length. tf-idf gains +0.61 when you paste the essay's own sentences back in
Presence detector — PERSUADEv0.1Role redefinedcounterclaim recall 0.036 at 85% precisionCannot grant encompassing credit. Useful only backwards, as an absence screen
Adversarial suite80 offline / 40 liveBuilt and run twicenonsense scores 2.60 vs 1.36 for weak writingFluent nonsense is still the hole. Through-line at 0.70 catches 4 of 6
Presence rubric + rater setv2, R1–R67Agent-tested onlyagent–agent κ 0.95–1.00That κ satisfies nothing. 8 gold labels need human adjudication before the queue can open
Shadow-trial dashboard24 seeded responsesLive, unrunneeds 5–10 teachers, then 50–100No teacher-appropriateness data exists
Session event stream & audit logNot built7 timing field groups, all newNone of the packet's timing requirements is derivable today
Experimental-arm machineryNot builtREQ-01, REQ-02, REQ-06The primary confirmatory test of the whole program currently has no engine.
Two config values are still wrong in production. SCORER_DAILY_LIMIT sits at 300 and belongs at 60 before students touch the system — flagged twice and closed neither time. THROUGH_LINE_MIN ships at 0.50 and the measured operating point is 0.70. Both are environment variables. Neither needs a deploy. This is the cheapest item on the entire map.
03

The research program

Three numbering systems run at once. They look like alternatives and are in fact layers: A2-3 is a coordinate — aim A2, bullet 3 — not a separate experiment. Read as layers, the program is compact.

Layer 1 · what we want to know

Program aims

A1 · A2 · B1 · C1–C5 · D1–D4 · E1–E5

Five themes (Ontology, Measurement, Instruction, Adaptation, Implementation) holding 16 aims, each with a numbered to-do list. “A2-3” means aim A2, bullet 3 — not a separate study.

Layer 2 · what we will test

Hypothesis registry

EDGE · SHAPE · GRAIN · TIME · MEAS · FEAS

21 hypotheses in six families, frozen at v3. Each names a prediction, a falsifier, and a sample size. One primary confirmatory test; no single result may edit the graph.

Layer 3 · what actually ran

Runs

runs/demo-2026-09-10, runs/a2-3-…

Six executed pieces of work, each with a ledger, verification counts and an approval record. This is the layer with findings in it.

What actually ran, and what it found

RunQuestionHeadline findingDid it change the build?
Cycle 1
37 agent runs, 9 stages
Does diagnostic-driven prerequisite sequencing beat a fixed sequence?No direct evidence exists either way in any located study. 119 claims from 77 publishers; 106 sourced, 81 independently verified, 9 disputed and corrected. H0 is the operative default; the burden is on H1.Yes — reframed cohort 1 from a confirmation study to a preregistered estimation cohort
A1-8
reciprocity stream
Is writing development hierarchical or reciprocal?Asymmetric-reciprocal, and it fades by grade 9. Transcription→quality β=.60 at grades 4–6 but .26 n.s. at 7–9. Grammar drill ES −0.32 vs strategy 0.82. 12 new claims, 12/12 verified.Yes — produced pathway v0.3: dropped CLAIM-01→PROC-01 and SENT-01→SENT-02, moved PROC-01 to level 4
A2-4
session-time audit
Is any node too narrow or too big for a 35-minute session?No node is wrong — the session is too short. All-12-ready by session 18: at 30 min 98/79/22%; at 35 min 100/99/60%; at 40 min 100/100/93%. Unplanned finding that mattered more: 12–20% of struggling students were being marked ready by chance.Yes, though indirectly — the false-ready finding is what produced unlock rule v1
B1-2
presence rubric
Can presence-with-meaning-preserved be coded reliably?Rubric v2 (67 rules) reached agent–agent κ 0.95–1.00 on five of six codes. The run states plainly that this is an upper bound on rubric clarity and satisfies nothing in MEAS-01. 12 disagreements on 8 items need human adjudication.Yes — three packet requirements, including the insertion flag as a database-generated column
A2-3
prompt stability
Is a placement a property of the student or of the passage?Designed, red-teamed, item-authored, readability-corrected, independently validated. Honest power: at n=70, INDETERMINATE is the modal outcome for true agreement 0.73–0.87. ~70–80 adult hours.No — and nobody is waiting for it. Explicitly on the do-not-build list, aimed at a registry v4 that does not exist
KLiCKE baselineWhat does honest human composition look like keystroke by keystroke?4,992 writers. Median 345 words in 29 min; 57% of the session producing nothing; pasting median 0. Process predicts quality only through fluency and length — so process is authorship evidence, not a scoring input. Your existing 40% paste rule sits above the 99th percentile of honest behaviour, and the flags do not fall harder on ELL writers (5.0% vs 5.7%).Yes — three shipped product changes
WAT instrumentShould a writing-apprehension measure go in?Daly-Miller 26-item adopted, verbatim from ERIC ED345265, ~5 min, before both cold prompts. KTEA-3 put on hold — off-construct for transfer, ~$1,147 for two forms, 30 proctored minutes each.Yes — decision dec_d4618a8448, packet REQ-20

The one test that decides everything

SHAPE-03 is the primary confirmatory hypothesis, and inside it sits the product thesis: does module-end mastery predict delayed unaided writing?

  • Needs ~190 analysed per arm. Cohort 1 at 30–80 students gives a minimum detectable effect of d ≈ 0.61–0.99 — an estimate, not a verdict, and the approval record says so.
  • The success metric was rebuilt after simulation showed the original had 12% power: counting mastered nodes correlates with delayed writing at r = 0.12 and needs N≈535. The replacement — mean success rate over the module's final third — correlates at r = 0.49, giving 80% power at N = 31, and correctly drops to 0.15 in a simulated “Quill world”.
  • The kill criterion is written down. If partial r < 0.10 with a CI upper bound below 0.20 across two cohorts, the componential thesis fails for this grade band and the product pivots to composition-first.

Work running with no consumer

Worth knowing before you fund more of it. None of this is wasted — but none of it is currently blocking or unblocking anything.

  • The prompt-stability study (A2-3). Fully designed and validated, then placed on the do-not-build list, aimed at a registry version that does not exist.
  • Edge annotations v2 and v3. All 14 edges still typed uncertain; the brief states human review of them blocks nothing in cohort 1.
  • The scorer research. The 80-item adversarial set, the bias tables and the evaluation harness have no live input while the calibration set does not exist.
  • The second research cycle. Seven stream briefs were scoped; only the reciprocity rerun was executed. Stream 1 — time-equated sequencing — is the one that would actually move the H1 evidence base, and it has not run.
04

The contract between research and build

The most unusual piece of the program, negotiated across four rounds. Research does not send prose to build. It emits a machine-assembled implementation packet; a building agent must validate it and refuse an unapproved one, then return a traceability report mapping every requirement to where it was implemented and tested.

The interface works because build pushes back. Across four rounds, research lost two of its three recommendations — one of them refuted on arithmetic rather than on evidence. That is the system functioning as designed, and it is the strongest single piece of evidence that this program is not an agent echo chamber.
RoundWho movedWhat happened
0Research → BuildCohort 1 does not depend on live model scoring — the primary outcome is adjudicated human double-rating; the scorer is shadow-only. Composition is a fixed 12-minute slot with an 8-minute floor and a hard cap, value to be decided by a human.
1Build → Research
four contradictions
Answered from the code, and corrected research on things that changed the arithmetic: the ready rule was 2-of-3 with retries counting, not 3-of-4 — so the false-ready figure was a floor, not an estimate. The engine computed a transfer-anchored verified state and gated nothing on it, which research called “the Quill failure mode written into the unlock logic”. Composition was already rung-based (3/6/9/14), not 12 minutes. And no arm concept existed at all. Build's closing line: “Nothing here is a decision. It arrived as a document from an agent run, not as a decision from Samuel Bennett.”
1bBuild → Research
the blocker
The two graphs were not the same object despite both having 13 nodes and 16 edges. The quotation dependency ran in opposite directions — the exact pair two registered hypotheses were designed to test. Verdict: nothing should be frozen until one is designated authoritative.
2Research → Build
concede three
Accepted all three corrections and re-ran the simulation under the real rule: 25% of node encounters marked ready while true success probability was below 0.70 (36% on SENT-02). Conclusion: “the rule is not slower, it is just wrong more often for weak students.” Three recommendations routed to you.
3Build decides one,
you decide two
Build refuted research's unlock recommendation on arithmetic: a disjunction can only unlock more often than its strictest branch, so “provisional OR first-try CR” scored 20.6% false unlocks — worse than the rule it replaced. Unlock rule v1 frozen instead at 12.9%. Your two decisions: dec_f353a96618 (variable slot, 30-XP goal, no cap at all) and dec_09f45341d4 (pathway v0.3 is the graph).
4Research re-derivesYoking moved from quantity to rule. Both arms must now run a byte-identical config, hashed per session and audited daily; composition minutes are expected to differ by arm and are reported rather than policed. The causal claim narrows from “sequencing at equal dose” to “the policy as deployed”. awaiting your sign-off
Packet v4

Superseded

Graph 0.2-reconstructed, 16 edges. 25 requirements, 24 telemetry rows. Wrong graph, wrong dose model, wrong mastery rule.

Packet v5

Superseded almost at once

The substantive rewrite: pathway v0.3, yoking by rule, the 30-XP session goal, unlock rule v1, and the CR bank gap as a P0 requirement. 27 requirements.

Packet v6

The live contract

28 requirements (19 are P0), 32 telemetry rows, 16 guardrails, 14 prohibited changes, 9 open human decisions, 14 items on do-not-build-yet. Only change from v5: the WAT.

05

What the program actually owns

Separate from the code: a set of licensed corpora, a set of measurements nobody else has run on them, and a task bank. The measurements are the underrated asset — several of them reverse things the founding brief asserts.

CorpusSizeLicenceWhat you may do with itState
ASAP 2.024,728 essays
grades 6–10
CC BY 4.0The only commercially trainable corpus. Carries ELL, economic, disability, race and grade labels on every essay — so it also carries all bias testingTrained · QWK 0.786
PERSUADE 2.084,440 spans
8,426 essays
CC BY 4.0 (Lab release)Argument-element detection. Licence reversal: the GitHub copy is BY-NC-SA; the Learning Agency's own release is CC BY — and that is the copy you hold. The brief's “pull from GitHub not Kaggle” rule is backwards hereShips as absence screen
KLiCKE4,992 writersCC BY 4.0The honest-human typing baseline behind every authorship flagBaseline computed
ELLIPSE8,890 essays
23 raters
CC BY-NC-SANothing trained on it may ship. ShareAlike is the clause that could reach into the product, not just the input. Measurement onlyMeasured, cannot ship
College Readiness Math434 review notes
321 items
CC BY-SA or CC BY (conflict)Mined into your 19-defect review taxonomy and the reviewer marginal-return curveMined
PIILO6,807 docsCC BY 4.0Would evaluate the PII scrubber upstream of the modelNot started
AIDE / Quest1,378 / 6,989CC BY 4.0AI-detection check; reading-comprehension taxonomyDeprioritised / shelved

Three measurements that change the brief

  • Two trained humans agree at about 0.5 on analytic traits, not 0.8. Across ELLIPSE's 23 raters: Overall 0.511, Grammar 0.463, Vocabulary 0.394. The founding brief's QWK ≥ 0.80 per-trait bar is not a bar anyone has cleared, including humans. The replacement: holistic keeps 0.80 against a two-rater average; per trait the bar becomes model–human agreement at or above human–human on the same essays.
  • The six traits are one thing measured six times. First eigenvalue 73.3% of variance; second 0.377. Expect organisation, sentence clarity and conventions to collapse into one language-control trait. The four argument traits are a different construct — and nothing public validates them, which is precisely why your own students are the only route.
  • Analytic rubrics defeat the length effect — that is the real case for traits. In ASAP, word count alone predicts the human score at ρ = 0.733. In ELLIPSE, trait scores run 0.093–0.304. Not a finer picture of sub-abilities; a scorer that stops rewarding length. That is the property to protect and the argument to make.

And one about the task bank

From 434 expert review notes on 321 AI-generated items — the closest public analogue to your own authoring pipeline:

  • 18.7% had a broken answer key or a duplicate correct answer, and that is a floor. “An AI-authored task bank is not a draft that needs polishing, it is a draft where roughly a fifth of the items are wrong on the thing that matters most.”
  • Two reviewers, not one and not three. New defects found by reviewer 1/2/3/4: 1.33, 0.97, 0.53, 0.25.
  • Reviewers barely agree on what is wrong. Mean overlap of defect sets 0.26; on “is the key broken”, 35% agreement. Budget human attention for skill alignment and model-answer correctness; automate rendering, house style, reading level and answer-position balance.

Applied to your 965: about 180 items with a broken key, and ~1,468 review passes to find them.

06

The critical path

Five things block the pilot. They are not equally urgent, because they are not all on the same track — and the one people reach for last is the one that is strictly serial and takes the longest.

TRACK A — STRICTLY SERIAL, LONGEST LEAD TIME Execute the DPA + published notice + consent process Collect cold paragraphs 100–150 real grade-9 scripts MEAS-01 rater study 2 trained humans, 2 rounds, κ ≥ 0.70 TRACK B — PARALLEL, CAN START TODAY Author ~280 CR items 5 nodes × ≥28 per course Review the 965 tasks ~1,468 passes, two reviewers each Build the arm machinery assignment, config hash, event stream Cohort 1 runs grade 9 · 30–80 students ~10 weeks The thesis is answered does module-end mastery predict delayed unaided writing? Track A is the schedule. Track B shortens with more reviewers, more authoring agents, more engineering hours. Track A is a queue of signatures and calendar weeks, and it cannot be compressed. Starting it late moves the pilot date by exactly that much.
Two tracks, one convergence. Track B is parallelisable — more reviewers, more authoring agents, more engineering hours all shorten it. Track A is not: the rater study needs real scripts, real scripts need consented students, and consented students need a signed agreement. Starting Track A late moves the pilot date by however long it is delayed.

The five blockers — click one to see what it releases

07

Decisions awaiting a named human

Agents wrote almost everything in this program. These are the items the design explicitly reserves for a person to decide, plus the commercial calls that go with them. Several have been open since the program began and are now load-bearing. Click a status to cycle it — the marks are shared with everyone who opens the page.

connecting…

Product and research

The XP schedule and post-goal policy

Named by the experiment designer as a blocker — every XP-derived threshold is parameterised on it. Under current values a rung-1 session reaches the 30-XP goal in 4–6 active minutes and a rung-4 session in 16–20. If students stop at goal, 18 sessions yield as little as 90 instructional minutes with a 53% assessment share. The 45/30/25 task mix cannot fit inside the goal.

Sign off yoking-by-rule — and the narrowed causal claim

The v7 re-derivation replaces equal-dose yoking with byte-identical config hashed per session. It also narrows what you can claim from “sequencing at equal dose” to “the policy as deployed”. Everything downstream — what FEAS-01 means, the fidelity thresholds, the preregistration hash — waits on this.

The per-hour denominator

Open from the start, and more load-bearing now that hours differ by arm. The registry says the ~28-minute diagnostic and verification minutes count with a sensitivity exclusion; the experiment's analysis plan excludes them. Idle threshold is 60 s in one document and 90 s in the other. Both are live.

Adjudicate the eight gold labels

S07, S16, S20, S21, S26, S30, S32, S36 — twelve disagreement records where both raters went against gold in the same direction. Held out of the rater-training queue until a named human rules. This blocks the training set, which blocks rater training, which blocks MEAS-01, which blocks everything.

Replace the per-trait QWK ≥ 0.80 gate

The founding brief still states it; ELLIPSE shows two trained humans reach ~0.5 on analytic traits; the W-Pal work declares it unreachable by these methods; the backend doc still cites it as the operative rule. Four documents, three positions. Proposed replacement: holistic keeps 0.80 against a two-rater average, per trait the bar becomes model–human ≥ human–human on the same essays, plus stability under attack.

Sign the metric spec and write the analysis plan

“Fix the primary outcome and analysis plan in writing before any student is enrolled” appears as a step in the explainer and as requirement S1 in the research program. It has no signature. Must include retention probes at 30, 60 and 90 days past the module — a one-month module's randomised intervals produce a median realised delay of about a day and a half.

Composition UI: planning affordance or not

plan_before_draft and goal_named need new UI, which changes the frozen composition structure — so it must be decided before the freeze, not after. Research's own position is that revision-event count and minutes-before-first-draft-text suffice for cohort 1.

The edge-review reporting format

An edge claim is three separable claims — dependency validity, instructional efficiency, diagnostic predictiveness — reported separately. No format exists, so EDGE results have nowhere to land in the change log. Related: the two-teacher edge review form is still outstanding, and its six flagged edges are supposed to replace the six the registry is currently frozen around.

Extension nodes — cap exposure or disclose the package

Students in the diagnostic arm finish earlier and get extension work. Either cap that exposure, or disclose that H1 is being tested as a package rather than as sequencing alone. None exist in the engine either way.

Does handwriting stay in the model path?

The answer sheet prints the student's code in a box and a photograph can contain a name the student wrote there. The PII scrubber cannot see inside an image. Either transcription becomes the redaction point, or handwriting comes out of the model path until it can be checked.

Commercial and legal

Which scholarship path

Closed: Team America Foundation cannot be an SGO as structured — it is a private operating foundation, and “operating” is a subtype, not an exemption. Converting it is explicitly advised against. Open: Path A (register VPA as an eligible provider — free, do it regardless), Path B (partner with ACE Scholarships — one phone call, but they set the award criteria), Path C (a Vail Valley community scholarship fund — only after Treasury finalises the disqualified-persons rules).

VPA's entity structure

Nonprofit status, bridge financing from private partners, a platform with investor-style economics, and a related scholarship organisation potentially funding your own students. These four interact, and the recommendation is one session with a nonprofit tax attorney covering all of them at once rather than four separate conversations.

Co-authorship or paid contract — decide before the first email

Crossley's lab may well reframe a paid-hours request as a research collaboration with co-authorship. Say “paid contract hours” explicitly in the first email if that is what you want. Your target hires are exactly the people who would want to write a paper about this — which is also why the contractor agreement needs an explicit IP assignment and a no-external-publication clause.

The consultant rate

Three benches, and conflating them is the failure mode. $30–40/hr for raters is genuinely generous — Measurement Incorporated pays $15, ETS about $25. But $30–50 is below market for the expert tier: budget $50–75/hr or flat fees of $1,500–5,000. The outreach templates still carry bracketed placeholders.

Price and bucket

$500–2,000 per student per semester fits inside ESA award sizes and inside a $1,700-granularity scholarship pool, and needs no purchase order, RFP or procurement cycle. Open question worth money: whether an accredited, credit-bearing course prices into a better-funded ESA bucket than “tutoring”.

One decision you already made that is worth restating. Never take advertising support. It permanently forecloses classroom use under the COPPA school-authorisation exception — which is the real mechanism behind Duolingo's schools exit, not the FERPA story. Worth correcting before that version reaches a pitch: an investor who knows the file will notice.
08

Outside parties, and what each is owed

Every one of these is drafted and none is sent. The outreach templates are written, the addresses are verified against live pages, the sequence is planned. The gap between this program and its next phase is largely a set of emails.

ASU Prep — open the FERPA conversation longest lead time on the list

FERPA's school-official exception runs through the school, not through you — you cannot self-designate contractors as school officials. De-identification handles rating; the formal studies exception is what you need to link a student's baseline essay to their delayed-transfer essay, which the success metric requires. Start it now; it will take longer than recruiting. Contacts: Cat, Jared, Meg.

Execute the VPA–Elevate Edwards data-processing agreement

Names the school-official designation, the authorised purpose, the ban on redisclosure and on training anyone's models, the retention periods and the deletion process. Until it exists, no student account should be cleared for consent. Alongside it: the published privacy notice carrying retention numbers verbatim, and the written parental-consent process.

Scott Crossley — LEAR Lab, Vanderbilt highest-value single email

He built ASAP 2.0, PERSUADE 2.0 and ELLIPSE — the three corpora your whole scoring plan rests on. The lab's jobs page explicitly offers paid student positions with no formal application; the lab is rebuilding after a move from Georgia State, so there is less competition for student hours. It also publishes on bias in LLM writing feedback. scott.crossley@vanderbilt.edu

Stacy Bailey — University of Northern Colorado

Best Colorado rater pipeline: she directs English Education and coordinates the MA, whose students are current and aspiring ninth-grade English teachers — the exact rater profile. Two hours from Edwards for in-person calibration. stacy.bailey@unco.edu

Peter Foltz — CU Boulder, NSF iSAT

Co-created the Intelligent Essay Assessor and ran large-scale automated scoring at Pearson. The most credentialed automated-essay-scoring person in Colorado, and has the industry instincts to understand a paid consulting ask without a grant wrapper. Aim him at adversarial robustness. peter.foltz@colorado.edu

Derek Briggs — CADRE, CU Boulder

A whole doctoral programme trained in rater agreement, validity, scaling and subgroup analysis; one email to the director reaches the cohort. Also there: William Penuel on research-practice partnerships. derek.briggs@colorado.edu

Joshua Wilson — University of Delaware

$3.5M+ in Gates, NSF and Spencer funding on automated writing evaluation efficacy. His research question is your governing risk. Small PI-centred group, fast personal replies. joshwils@udel.edu

Danielle McNamara — broker through ASU Prep, do not cold-email

She built Writing Pal, the closest historical analogue to this product including its own transfer lessons. You have already rebuilt 28 of its indices and measured where they fail. Also at ASU: Maria Goldshtein, an “Inclusive Language Analytics” postdoc — a near-exact match for the subgroup-fairness reporting requirement.

Kyle Cureau — take the coffee

Four things to get: the economics of a fine-tuned classifier versus LLM calls at your volume and latency (your open scorer-architecture question, and he has worked it for a consumer product at scale); whether Barak Oshri or Violet X. would review the trait-scoring design — you have no ML advisor and QWK-with-subgroup-reporting is a real ML problem; what he learned from 100+ interviews about why people don't write; and the channel conversation. Be explicit that his success metric is engagement and yours can never be — an investor will conflate the two.

ACE Scholarships — one phone call, and it has a date

Denver-based, already operating in Colorado, explicitly markets SGO infrastructure. Foundation supporters donate and take the credit; VPA families apply. No entity to form, no 90% rule, no public support test, no annual state certification. The trade is that ACE decides who gets funded. Must happen in 2026, before the 1 January 2027 start.

Report the SGO finding back to the foundation

Team America Foundation cannot become an SGO as structured, and converting it is an expensive way to get a worse version of the ACE partnership. This needs to go back to the foundation as a finding, not sit in a document. Also confirm whether TAF files Form 990-PF — that is the definitive confirmation of private-foundation status. TAF could donate to a future SGO, but a private foundation granting to an organisation it effectively controls raises self-dealing questions.

Counsel — two engagements, both single sessions

Nonprofit tax: entity structure, bridge financing, investor-style platform economics and a related scholarship org, covered together. Privacy and education: review the schema and privacy document, rule on whether Colorado's student-data act adds anything beyond FERPA and COPPA, review the contractor agreement template once so it can be reused, and confirm the corpus licences before anything ships.

Anthropic data terms in writing

No training on commercial API data by default, but zero-retention eligibility is unconfirmed. Get it in writing and attach it to the DPA as the subprocessor record, before the pilot.

Two teachers — return the edge review form

Still outstanding. Six edges are flagged suspect and undecided, and the registry's six EDGE hypotheses were chosen by graph topology rather than by anyone who teaches ninth grade. The form exists and is ready to hand over.

5–10 teachers for the shadow trial

The tool is live with 24 seeded responses: blind rubric scoring, then reveal, then helps / should-defer / show-feedback verdicts, with teacher–model and teacher–teacher agreement computed automatically. Expands to 50–100. Produces the only teacher-appropriateness data that will exist.

Draft the one contractor agreement

FERPA-aware confidentiality, explicit IP assignment, no-external-publication, 1099 classification safeguards. Counsel reviews it once; reuse forever. Watch three things that will bite: F-1 candidates need written DSO confirmation before signing, funded RAs face outside-work caps, and W-9 before first payment.

The Learning Agency — licence terms

Steward of ASAP 2.0 and PERSUADE going forward, and the source of the licence discrepancy you found. Worth a conversation regardless. Perpetual Baffour there works on bias in automated essay scoring. info@the-learning-agency.com

09

The cleanup queue

Work at this pace leaves a wake. These are places where two project documents disagree, or where a file says something that is no longer true. None is fatal; several would not survive an accreditor’s read, and two would corrupt a study if left alone.

What disagreesThe two versionsWhy it matters
The rater training set contradicts itselfNine items carry v2 gold labels with v1 rationales still attached, and S30's rationale cites a rule that v2 retiredFix first This file seeds the rater training queue. Trainees would read reasoning that contradicts the labels they are scored against
Two Gettysburg transcriptionsThe rater set ships the Gutenberg text (“cannot”, “from this earth”); the rubric anchors quote the Bliss forms (“we can not”, “from the earth”)Fix first A rater checking a quotation against the passage copy will hit a false mismatch on a rubric about quotation accuracy
Three different ASAP headline numbersQWK 0.807 (n=7,421 official split) vs 0.786 (80/20 by essay). Length baselines 0.724 / 0.668 / 0.652The readiness gate is stated as “beat 0.786”. Which split is canonical is not settled anywhere, and the difference decides whether the scorer passes
Subgroup bias points two directionsOne document reports signed bias of −0.110 for ELL and −0.149 for disability and concludes lower-scoring groups get pushed further down. Another reports mean signed error within 0.05 of zero for every subgroup and concludes the gaps are in agreement, not in levelDifferent models on different splits — but the documents do not say so, and only one of the two can be quoted to an accreditor
Seven traits or eightThe rubric spec says eight traits with 0–4 anchors; the grading tool scores seven; the training doc says sevenNothing reconciles it. The missing one appears to be independence, which is human-only — but that should be written down, not inferred
The experiment is still written for grades 6–8Grade 9 is decided; the experiment's population, moderators, strata and inclusion criteria still read 6–8, and the MDE figures still need correcting to 0.61 / 0.77 / 0.99The preregistration hash cannot be taken over inconsistent text. Also pending: one claim needs relabelling as unverified, and an instrumentation field still describes the retired fixed slot
Pathway size is stated four ways188 nodes (brief) · 203 (deployed at /graph) · 34 (engine v0.1) · 13 (what actually runs)Nothing maps the 13 pilot nodes onto the 188, and the brief calls for a 25–35-node induced subgraph. The 34-node version seems abandoned without being declared so
Registry text describes a retired designFEAS-01 still says “8-minute floor / hard cap”; standing rule 4 still says composition minutes are yoked; the cohort scope still says “feasibility of equalizing writing time”All three were superseded by a later decision. They need a human-approved amendment, not a quiet edit
Approved files still say “proposed”Every change record inside the approved pathway file still reads proposed (designed, not decided) though the file-level status is approved and frozenMinor, but it is the provenance trail an accreditor reads
The site map is eighteen bundles staleIt points /writing-app at bundle -2; the current one is -20. No document states which bundle is live on Squarespace right nowIt is also the only document describing the review page, the 203-node graph and the legacy diagnostic
Task review counts disagree965 pending in one document, 734 in another; 232 legacy tasks versus an implied 231Small, but it is the denominator of your largest content job
Corpus sizes disagree with the briefASAP 24,278 (brief) vs 24,728 (held); ELLIPSE “~6,500” (brief) vs 8,890 collected / 6,468 published — and the brief groups it with the commercially usable sets when it cannot shipThe ELLIPSE one is a licensing error in the founding document, not a typo
Threads with no owner. Across all 52 documents, four items appear in the record with no status and no next step: Cognia is named as part of the credit tier and has no status anywhere; the NCAA item surfaced only as a passing digest note that the credit tier must be legibly not credit recovery; the exit survey is referenced as needed and has never been drafted; and the question of whether the app's pathway view should link out to the full graph has been open since the site was first mapped. Also unresolved: the 12 GB of licensed corpora live in a laptop’s Downloads folder, and ELLIPSE exists only in a cloud workspace.
10

The outside view

Everything above is the operator’s view. This is the version for someone coming to it cold: what is real, what is still a bet, what the next round of resourcing buys, and what would make us stop.

What is real

A working platform and a research organisation

  • A deployed student app running a 13-skill pathway over 965 authored tasks, with placement, adaptive selection, spaced review, a composition every session, and a decision log that records a written reason for every choice it makes.
  • A scoring function live in production that resists the attacks that beat the statistical baseline — padding an essay by 40% costs it a third of a point per trait, where the conventional model gains 0.61.
  • A backend with consent enforced in the database rather than the interface, retention enforced by code, and a teacher grading tool that computes agreement automatically.
  • A research program with 21 registered hypotheses, a frozen registry, a written kill criterion, and a machine-readable contract that a building agent must validate before it implements anything.
What is still a bet

All of it, honestly — and that is the design

  • No student has used this. Every number on this page comes from public corpora, simulations or agent runs. The simulations are good — they caught that the original design produced zero compositions in twenty sessions, and that the original success metric had 12% power — but simulated students are not evidence about real students, and the documents say so on every page.
  • Nobody has shown that prerequisite sequencing works for writing. The literature cycle's conclusion was that no direct evidence exists in either direction, so the default is that it doesn't, and the burden is on us. That is an unusually honest place to start and it is why cohort 1 is framed as an estimation cohort rather than a confirmation.
  • The scorer has never been checked against a human. Not because it failed — because the measurement has not been run.

What the next round of resourcing buys

  • A rater bench. Two trained humans on 100–150 real ninth-grade paragraphs at $30–40/hr, plus one or two expert consultants at $50–75/hr. This single item unlocks every measurable claim in the program, and it is the difference between “the model agrees with itself” and “the model agrees with teachers”.
  • Content review hours. ~1,468 review passes over 965 tasks, where the public evidence says one item in five has a broken key.
  • ~280 new constructed-response items so that five of the thirteen skills can be gated on writing rather than on selection.
  • Engineering for the experimental arms — currently the primary test of the whole thesis has no engine.
  • Two single-session legal engagements, one tax and one privacy.

What would make us stop

This is written down and it is unusual to have:

If module-end component mastery does not predict delayed unaided writing — partial r below 0.10 with a confidence interval topping out under 0.20, across two cohorts — then the componential thesis fails for this grade band and the product pivots to composition-first with strategy instruction.

The whole design exists to get that answer in month three rather than year three. The documented failure of this product category is transfer — drill gains that never reach real writing. Quill's own randomised trial moved a sentence-combining test by d = 0.30 and moved actual composition by nothing distinguishable from zero. Every structural decision here — composition from session one, mastery defined as predicted use in free writing, more writing rather than more drill when the two diverge — exists to avoid rebuilding that product.

Where this stands. The build is further along than the evidence, and the evidence is further along than the paperwork. That is the correct order to have built things in — but it means the next phase is not more building. It is two signatures, two raters, and a hiring round.