What researchers studied
Nationwide quasi-experimental study using administrative data on Chilean primary-school teachers and students from 2005 to 2015. A 2011 reform required teachers rated 'basic' to be re-evaluated after two years rather than four, enabling difference-in-differences analysis.. Nationwide administrative data on Chilean public primary-school teachers and students observed from 2005 to 2015; the public abstract does not provide one consolidated participant count
What they found
- The reform increased the probability of two-year re-evaluation by 58.9 percentage points for affected lower-performing teachers.
- In descriptive pre-post comparisons, teachers who completed re-evaluation improved their evaluation scores by about one standard deviation.
- Difference-in-differences estimates ruled out student-achievement effects larger than 0.04 standard deviations.
- Researchers found no meaningful changes in student-reported teaching practices or caregiver-reported teacher caring, cautioning against treating large evaluation-score gains as sufficient evidence of improved teacher effectiveness.
What the study does not prove
- This is a 2026 EdWorkingPaper and should not be described as a peer-reviewed journal article unless a later publication is independently verified.
- The study evaluates a specific national teacher re-evaluation policy in Chile rather than private tutoring, U.S. teacher licensure, or Noor Lyra's educator-selection system.
- The roughly one-standard-deviation improvement in teacher evaluation scores comes from a descriptive pre-post comparison, whereas the student-outcome conclusions rely on a quasi-experimental difference-in-differences design; these should not be conflated.
- The findings show that more frequent formative evaluation alone did not generate detectable student-learning improvements of meaningful size in this setting; they do not show that evaluation, coaching, or professional development are generally ineffective.
Evidence strength: Promising quasi-experimental evidence using nationwide administrative data; difference-in-differences design; EdWorkingPaper.
Why this matters for families
A strong educator rating can be useful information, but it is not the same thing as proof that students are learning more. This 2026 Chilean study found large improvements in teacher evaluation scores after re-evaluation without comparable measurable gains in student achievement.
Noor interpretation
How Noor translates the evidence into practice
The Noor-relevant lesson is that an educator can become better at the metric being evaluated without necessarily producing better student outcomes. Internal quality systems should therefore avoid over-relying on credentials, observations, or scorecards alone. Educator development is strongest when evaluation is connected to evidence of what students understand, retain, and can do independently.
For Noor Lyra Educators, observations and performance reviews should be paired with student evidence: baseline work, misconception patterns, progress over time, parent-visible outcomes, and independent student performance. Coaching should target instructional behavior that is plausibly connected to learning rather than optimizing merely for an evaluation rubric.
Read the original source
Noor links to the original or authoritative source so families can distinguish the evidence itself from our interpretation.
Open original source →DOI: 10.26300/ms3w-2s74
Research notes
No single study determines a student's plan. Noor uses research as one input alongside the learner's goals, observed performance, academic context, and response to instruction.