Writeiq's marking framework, IWAF, is built on decades of writing-assessment research and is benchmarked at engine level against expert human scores on openly licensed external corpora. This page sets out the methodology, the theoretical underpinnings, the evidence with its limits, and the work in progress to publish the framework formally.
IWAF is a seven-dimension rubric for assessing extended writing across Years 3 to 12, developed by trained assessors and leading educators across Australia's government and independent school sectors. The dimensions are Voice, Structure, Cohesion, Vocabulary, Sentence Craft, Text Structure, and Conventions. Each is scored on a four-band scale - Emerging, Developing, Consolidating, Extending. Writing is judged against the stable Secondary framework; the approved year-level interpretation is applied afterwards. That is how a single methodology can answer "is this a strong piece for a Year 4 writer" and "is this a strong piece for a Year 11 writer" in the same vocabulary.
IWAF 3.0 (the current version) sets band thresholds at 0.31 / 0.56 / 0.80 against the maximum criterion score. These are current design thresholds, set in April 2026 to align with the MTSS tier distributions documented by AERO, Fuchs (2010) and the National Center on Intensive Intervention. The evidence behind each boundary differs, and we say so: the 80% Extending threshold has strong grounding in Bloom’s mastery learning framework (1968) and Rosenshine (2012); the 56% boundary has moderate support in the standards-referenced assessment literature; the 31% boundary is a design decision we are validating through our own calibration programme, and we do not describe it as research-derived. The full evidence breakdown, threshold by threshold →
IWAF is not an opinion about writing. It is a synthesis of established assessment and instructional research, made operable for everyday classroom use.
The 0.80 threshold for Extending follows mastery-learning research from Bloom (1968) onward, refined through Guskey's school-implementation work and the Australian Education Research Organisation (AERO) mastery framing. 80% is where consistent, transferable competence begins.
The Text Structure dimension draws on the Sydney School genre tradition (Martin, Rothery, Christie) and Systemic Functional Linguistics. Different writing purposes have different shapes; the rubric reflects this rather than treating "structure" as a single skill.
Hattie & Timberley's "Where am I going / How am I going / Where to next" framing shapes the per-piece feedback structure: each marked piece returns a band, evidence, and one explicit next-step focus. Feedback is forward-looking by design.
Pearson & Gallagher's GRR model (1983, refined since) underpins the auto-generated lesson plans: I do, We do, You do. Every marked submission produces a lesson plan ready to teach against the patterns the cohort actually showed.
Rosenshine's Principles of Instruction shape the way the marker writes growth feedback: small, specific next steps the student can practise and the teacher can model. No generic "develop your ideas further" advice.
A full reference list accompanies the published technical white paper. The references below are the load-bearing ones.
A marker is only as good as its agreement with independent human judgement, so Writeiq's engine is benchmarked against expert human scores on the openly licensed PERSUADE and ASAP corpora: large, independently assessed collections of student writing. Held-out testing was strict, so the data used to calibrate was never the data used to judge. This is engine-level evidence against holistic external rubrics; it is not a year-level accuracy claim.
Across those tests, the spread of Writeiq's scores tracked the human markers rather than collapsing toward a safe middle band. Distribution and tails are reported alongside agreement, because a single number can hide both.
We publish engine-level benchmark evidence against external rubrics. We do not publish a year-level accuracy claim, and no year level from 3 to 12 carries one.
On criterion-level judgements, pairs of trained human raters agree exactly only about half the time, a finding from the published inter-rater research: exact agreement has a human ceiling. Our published target is agreement within the range two experienced teachers reach with each other; a blind moderation programme measures progress toward it.
The year-level interpretation layer is versioned and sealed; it changes only through a governed release process, and official results are never silently recalculated. Alongside it, the ongoing Writeiq calibration programme keeps collecting in-school evidence: teachers run blind calibration exercises inside the product and record fairness judgements on aligned results, and that evidence stream gates any expansion of the interpretation's scope.
Feedback is validated separately rather than assumed to follow from the scores. In product, every feedback phrase passed line-by-line educator review, and feedback must quote the student's own writing as evidence. The blind moderation programme measures the quality of Writeiq's suggestions against teacher judgement as the evidence base grows.
Qualifications, because they matter: the benchmark figures are engine-level results against holistic external rubrics on openly licensed corpora; they are not an Australian year-level calibration. Writeiq is not affiliated with or endorsed by any national assessment body, and Writeiq results are not official results in any external assessment program. Teacher acceptance of Writeiq's suggestions is a different measurement from external accuracy, and we never blend the two into one number. The full methodology, confidence intervals and limitations are in the technical white paper.
An honest research page sets out what an instrument can do and what it cannot. The list below applies to Writeiq as it ships in 2026.
The full technical white paper is published and free to download: how IWAF was constructed, the deterministic rule system (including the rules that were removed as unvalidated), the year-blind judgement and the versioned interpretation layer, production release governance, the international benchmarks, and what Writeiq deliberately does not claim. It is written to be read by school leaders and audited by psychometricians.
Download: Marking Once, Honestly - The Measurement Architecture of Writeiq (PDF, 14 pages)
The next study is pre-registered. The Multi-School Blind Agreement Study (MSBA-1) has its thresholds, sample rules, exclusions and analysis plan frozen before any data exists, so when the results are published you can verify the bar was set in advance and never moved. The document cannot be edited: our build system fails if it changes, and its fingerprint is published here so any copy can be checked against it. Because it is frozen, its background section describes our evidence base as at the date of registration; the current evidence position is always this page and the latest edition of the white paper.
Download: MSBA-1 pre-registration (PDF, 4 pages)
SHA-256 f6ab1df172951de359430f14b3971da84d53a6baadf75b3d29fc47267c46bee9
If you'd like to see the underlying registers and test suites, discuss anonymised trial data, or explore a research collaboration, we're happy to share under a confidentiality agreement.
hello@edsthetic.com.au · reference "IWAF research enquiry" in the subject line and we'll get back to you within two business days.