ILDF 2.0 Entry Stack · Step 7 · Phase 3 — Evaluation: Local Impact

Determine Your Formative Evaluation Criteria Before the Pilot.

Formative Evaluation Designer builds one formative evaluation plan for one pilot of one design — Competency, Evidence, Task, and Assembly, in that order, so what counts as “it worked” is committed before any pilot data exists to argue about.

What this process is grounded in ↓

13 sources behind this tool's steps and structure, from Guerra-López's Impact Evaluation Process to Google's PAIR Guidebook — click to see them all.

  • Overall process, walked step by step across this tool: Guerra-López's Impact Evaluation Process (2007/2008) — stakeholders and objectives, measurable indicators, data sources and collection methods, data analysis, communicating results.
  • Competency → Evidence → Task → Assembly structure: Evidence-Centered Design, Mislevy, Steinberg & Almond (2003), “On the Structure of Educational Assessments,” Measurement: Interdisciplinary Research and Perspectives, 1(1), 3–62.
  • Stealth/embedded assessment lineage: Shute (2011); Shute (2023), “The History of Stealth Assessment and a Peek Into Its Future,” in Games as Stealth Assessments, IGI Global.
  • Two-track measurement (performance + SRL/process indicators), one quasi-experimental study, not industry consensus: Xu (2025).
  • Formative-evaluation continuum: Tessmer (1993); Dick & Carey (1990).
  • Thematic analysis, one example way of reviewing collected evidence: Braun & Clarke (2006); applied version: Rosala (2022), Nielsen Norman Group.
  • Success Case Method: Brinkerhoff (2003).
  • Argumentative-grammar audit (applied to real findings in Learning Analytics Interpreter, Entry Stack Step 8, not in this tool): Middleton, Gorard, Taylor & Bannan-Ritland (2008).
  • Pre-commitment gate against post-hoc scope creep: Hardman (2026).
  • Current field validation, GenAI + stealth assessment: Rahimi, Shute & Almond (2026), guest editors' preface, Journal of Research on Technology in Education, 58(1).
  • Success-metric threshold framework, Wizard-of-Oz testing, and qualitative-feedback-in-early-phases practice: Google's People + AI Research (PAIR) Guidebook, pair.withgoogle.com, updated for the generative-AI era; industry-standard practitioner resource, not a single study.
  • Structured need-finding methods, referenced from PAIR: IDEO Design Kit, designkit.org/methods.
  • Deciding whether an AI-assisted tool is actually worth adopting: the PROVE framework, Nielsen Norman Group (2026).
1Competency ModelNot started
2Evidence ModelNot started
3Task ModelNot started
4Assembly ModelNot started
Competency Model

What is this pilot actually testing?

Name the specific competency claim(s) the design is meant to support — distinct from tool-use fluency. This step forces the real target construct into view before anything else gets built.

Why start here, before anything else: this tool's four steps (Competency → Evidence → Task → Assembly) are Evidence-Centered Design applied to a real pilot (Mislevy, Steinberg & Almond, 2003) — you name the claim first, then decide what would actually count as evidence for it, before any pilot data exists to argue about. It's the same design logic behind stealth assessment: assessment gets designed into the experience from the start, not added in after development wraps (Shute, 2011, 2023). Skip this order and the most common failure follows: what counts as “it worked” quietly reshapes itself around whatever data you happen to get.

Nothing you enter is stored on our servers. Your text is processed by Anthropic's Claude to generate this tool's response. See Privacy for details.

Regenerations used: 0/3
Your committed claims

No claims committed yet — add one from the candidates above, or type your own below.