In most post-Soviet and transitional systems, a reform lives or dies by a single number: the average score on a national or international test. Ministries quote it. Donors demand it. School inspectors turn it into a traffic-light chart and move on. But a test score is a thin slice of a reform that is supposed to change what teachers do in classrooms, which languages children learn in, and how local curricula get made. This article lays out a wider evaluation frame for classroom-level implementation of language-of-instruction policy, curriculum localization, and teacher professional identity. It is written for practitioners, teacher educators, and policy analysts working in Kazakhstan, Uzbekistan, Kyrgyzstan, Tajikistan, Georgia, Armenia, Azerbaijan, and Moldova, where the distance between policy text and classroom practice is often the real reform story.
I use the term classroom-level implementation to mean the observable changes in lesson structure, language use, materials, and teacher decision-making that follow a reform. Adjacent concepts include fidelity of implementation, curriculum localization, language-of-instruction shift, and teacher professional identity. These are not the same as test-score gains. A school can raise its average score by narrowing the curriculum, excluding children with weaker language skills, or teaching to a past paper. A school can also improve classroom practice without an immediate score change because teachers are learning new routines, parents are adjusting to a new language policy, or textbooks have not yet arrived. Evaluation that ignores these processes will misread both failure and success.

Why Test Scores Are a Weak First Measure
Test scores are attractive because they are cheap to collect, easy to compare, and familiar to international donors. But in the systems I study, they carry three specific problems.
1. Score inflation through curriculum narrowing
When a reform is judged mainly by a test, schools respond by teaching what is tested. In several Central Asian and South Caucasus systems, this has meant reducing hours for local literature, history, and oral language work in favor of drill in mathematics and the dominant language of assessment. The reform may have intended to strengthen local curriculum, but the evaluation incentive pushes the opposite direction. A score rise can therefore indicate a loss of the very content the reform was meant to protect.
2. Exclusion and selection effects
Average scores can improve because weaker students leave the sample. In systems with high dropout, migration, or informal exclusion of children with disabilities, a rising average may reflect who is still sitting the test rather than what teachers are doing better. Language-of-instruction reforms are especially vulnerable here: if a school shifts from Russian to Kazakh or from Russian to Georgian, some families move children to another school or another country. The remaining cohort may be more advantaged, and the score may rise for reasons unrelated to teaching quality.
3. Time lag and threshold effects
Classroom change takes years. A new language-of-instruction policy may require teachers to relearn subject vocabulary, rewrite lesson plans, and build new assessment routines. During that period, scores often dip before they rise. If evaluation stops at the test, the reform is declared a failure at exactly the moment when implementation is beginning to stabilize. A materialist evaluation has to account for the cost of transition, not just the endpoint.
A Broader Evaluation Frame: Four Layers
I propose a four-layer frame for evaluating reform in transitional education systems. Each layer answers a different question and requires different evidence.
Layer 1: Policy-to-classroom translation
This layer asks: What did the reform actually ask teachers to do, and what did they understand it to mean? Evidence includes policy documents, curriculum guides, teacher manuals, and interviews with teachers about what they think the reform requires. In Kyrgyzstan, for example, a multilingual education policy may be read by one teacher as “use Kyrgyz for all subjects” and by another as “use Kyrgyz for greetings and switch to Russian for content.” Both readings are implementation, but they are different reforms. Evaluation must capture that variation before it can judge outcomes.
Layer 2: Classroom routines and language use
This layer asks: What changed in the actual lesson? Evidence includes structured classroom observation, audio or video recordings, and teacher logs. Observers should record the language of instruction for each activity, the type of student talk, the use of local texts, and the time spent on different tasks. A reform that is supposed to increase student talk but produces 40 minutes of teacher monologue has not been implemented, whatever the test score says.

Layer 3: Teacher professional identity and workload
This layer asks: How did the reform change what teachers think they are for, and what did it cost them? Evidence includes teacher interviews, time-use diaries, and analysis of salary and contract changes. In many post-Soviet systems, teachers are asked to become curriculum developers, language assessors, and community liaisons without additional pay or reduced teaching hours. A reform that succeeds on paper but burns out its teachers will not last. Evaluation should measure workload, attrition, and the shift in professional identity from “transmitter of the state curriculum” to “local curriculum maker.”
Layer 4: Material conditions and funding
This layer asks: What did the reform assume about buildings, books, heating, and salaries, and what was actually provided? Evidence includes budget analysis, textbook distribution records, and school infrastructure audits. A curriculum localization reform that assumes every teacher can print local materials is not the same reform in a school with one working computer and no paper budget. Political economy matters here: who pays for the reform, who profits from textbook contracts, and who absorbs the unpaid labor of implementation.
Tools and Methods That Fit the Region
International evaluation frameworks often assume stable schools, reliable data systems, and a research budget. In the systems I work in, evaluation has to be lighter, cheaper, and more embedded in existing routines. Three methods are practical.
Structured classroom observation with a short rubric
A 10-item rubric can capture language of instruction, student talk, use of local materials, and lesson structure. Observers need training and a moderation process, but the tool itself can fit on one page. In Kazakhstan and Georgia, school inspectors already conduct lesson observations; the task is to align their rubrics with the reform’s actual goals rather than with a generic checklist of teacher behaviors.
Teacher implementation logs
Teachers record, once a week, what they taught, in which language, using which materials, and what they changed from the previous week. Logs are self-report and therefore imperfect, but they capture variation across schools and over time. They also give teachers a voice in the evaluation, which matters for professional identity. The cost is low: a paper form or a simple spreadsheet.
Document and textbook analysis
Curriculum localization can be evaluated by comparing the official curriculum, the approved textbook, and the materials teachers actually use. In Uzbekistan and Tajikistan, textbooks often arrive late or not at all, and teachers improvise from Soviet-era materials. A document analysis that maps the gap between policy, textbook, and classroom material is a direct measure of implementation depth.
What to Do with the Evidence
Evaluation is not an academic exercise. It should feed back into policy, teacher education, and funding decisions. Three uses are most relevant to this region.
1. Adjust the reform before it fails
If observation shows that teachers are using the new language for greetings but not for content, the ministry can adjust teacher training and materials. If logs show that teachers are spending three extra hours a week on reform paperwork, the ministry can simplify reporting. Early evidence is more useful than a final score.
2. Defend the reform against premature cancellation
Reforms in transitional systems are often cancelled after one political cycle. A broader evaluation can show that implementation is progressing even when scores have not yet moved. That evidence can protect a reform from being replaced by the next donor-driven project.
3. Name the funding gap
When evaluation documents that teachers are buying their own paper, heating is off in January, or textbooks arrived in April, it makes the funding constraint visible. That is a political act, but it is also a factual one. A materialist evaluation does not pretend that reform is only about pedagogy.

Common Mistakes in Reform Evaluation
I see the same mistakes repeated across the region. They are avoidable.
Mistake 1: Evaluating the policy text, not the classroom
A ministry may produce an excellent curriculum framework and then declare the reform complete. The framework is a plan, not an implementation. Evaluation must follow the policy into the school, the timetable, and the lesson.
Mistake 2: Using only international test data
PISA, TIMSS, and PIRLS are useful for cross-country comparison, but they are not designed to evaluate a specific language-of-instruction or curriculum localization reform. They are also administered in a limited set of languages and may exclude exactly the schools most affected by the reform.
Mistake 3: Ignoring teacher time
Every reform adds tasks: new lesson plans, new assessments, new reporting. If evaluation does not measure teacher time, it will miss the main reason reforms stall. Teacher workload is not a soft issue; it is a material constraint on implementation.
FAQ: Evaluating Reform Beyond Test Scores
What is the best single indicator of classroom-level implementation?
There is no single best indicator, but the most informative is usually the language of instruction during content activities, observed directly in a sample of lessons. If a reform is about language policy, that indicator tells you whether the policy has reached the classroom. If the reform is about curriculum localization, the best single indicator is the proportion of lesson time using locally produced materials rather than imported or Soviet-era textbooks.
How can schools evaluate reform without a research budget?
Use existing structures: school inspectors, methodologists, and teacher professional development days. A short observation rubric and a weekly teacher log cost almost nothing. The key is to agree on a small set of indicators, train a few people to collect them consistently, and review the evidence every term. You do not need a university research team to know whether teachers are using the new language or the new materials.
Why do test scores sometimes fall after a good reform?
Because implementation has a transition cost. Teachers are learning new routines, students are adjusting to a new language of instruction, and materials may not yet match the new curriculum. During that period, scores often dip. If the reform is well designed and funded, scores typically recover after two to four years. A fall in the first year is not, by itself, evidence of failure.
What role should teachers have in evaluating reform?
Teachers should be co-evaluators, not just subjects of evaluation. Their logs, interviews, and professional judgments are evidence. More importantly, when teachers help define what counts as good implementation, they are more likely to treat the reform as their own work rather than as an external demand. That shift in professional identity is itself a reform outcome worth measuring.
A Practitioner-Facing Provocation
Here is the question I leave with school directors, teacher educators, and ministry staff: If your reform were succeeding in every classroom but the test scores had not yet moved, would you be able to prove it? If the answer is no, then your evaluation system is measuring the wrong thing. Build the observation rubric, the teacher log, and the document analysis before the next reform arrives. Otherwise, you will be left arguing with a number that cannot see the classroom.
This article is part of a series on classroom-level implementation in post-Soviet and transitional education systems. A follow-up piece will examine how teacher professional identity shifts during language-of-instruction reforms in Kazakhstan and Georgia, with a focus on the unpaid labor of curriculum translation.






