Showing posts with label Value-added measures. Show all posts
Showing posts with label Value-added measures. Show all posts

Friday, January 2, 2015

Standardized Testing: The Final Frontier

Tests seem so reasonable at first --  teachers teach, students learn, and demonstrate mastery by passing a test. But as Daniel Koretz says at the start of his 2008 book, Measuring Up: What Educational Testing Really Tells Us, “Achievement testing is a very complex enterprise, and as a result, test scores are widely misunderstood and misused.” Now that is what I call an understatement. Furthermore, despite Common Core claims that better standards and tests mean fewer reasons for concern about their misuse, as Vito Perrone of Harvard University pointed out, “Most items on these various standardized tests remain well within the longstanding technology of testing, primarily to support the mechanical scoring procedures. They still seem to be limited instruments with too much influence” (1999, p. 152).

The testing “enterprise” is poised for a warp drive record-breaker of misuse insanity. In a nutshell, here’s how they plan to connect the dots.

A tiny fraction of what a student knows and can do is hypothetically captured, with some modicum of so-called scientific accuracy, by converting the number of correct answers out of the total number of questions on a standardized test to a raw score. Keep in mind that this single raw test score is still prone to error in its intent to measure what the student knows as the student may have made random choices, guessing correctly (or not), or may simply have had other contextual reasons for the performance including illness, distraction, nerves, etc. The test is also imperfect by design and is likely biased in some ways.

Now that raw score goes through some psychometric process to either be normed to a scale comparing it to other test scores, and/or it is ranked somewhere between unacceptable and excellent based on someone’s judgment of what students should know and be able to do. This is where all hell breaks loose as that converted score gets used.

How might it get used? For one, to tell the students and the parents or guardians how “well” they did which can involve labeling the converted score with a percentile rank, a grade-level equivalent, or just a descriptive meaning such as “meets standard.” However, it will likely be used in what is called a “high stakes” way to assign students to special education, to hold them back a year, or to track them into homogenous groups.

The most pernicious use is to group the scores to make claims about the quality of individual teachers. From there, it’s easy to see how tempting it is to make a claim about the quality of a school, and then a whole district. While we’re at it, let’s compare counties, states, regions, countries.

The cold hard truth, in Koretz’s words, is this:
Scores on a single test are now routinely used as if they were a comprehensive summary of what students know or what schools produce (p. 44-45).
He goes on later to add:
Simply attributing differences in scores to school quality or, similarly, simply assuming that scores themselves are sufficient to reveal educational effectiveness, is unrealistic. And more generally, simple explanations of performance differences are usually naïve. All of this is established science (p. 142).

Things get really tricky when hierarchical linear modeling kicks in to provide a “value-added” way to compare actual scores to a prediction and to use the difference to rate teachers’ effectiveness. Ignoring warnings from experts, these value-added models or VAMS, have been misused by policymakers to weigh heavily in the annual evaluation of teachers. Carol Burris, an outspoken principal who opposed this misuse of standardized test scores, recently wrote of a teacher’s lawsuit filed in New York State by my friend, Sheri Lederman, who hopes her case can become “a tipping point” in bringing this damaging unreliable practice to a grinding halt.

That may be wishful thinking because now the dots are being connected to the colleges and universities that educate teachers. They too are to be evaluated and ranked based on the performance of their candidates for teacher certification on standardized tests, which can be more than four in some cases. New federal regulations currently open for public comment until February 2nd would require these institutions of higher education to also track their teacher graduates, and collect their annual evaluation ratings including the VAMS measure, in order to be considered eligible for the TEACH grant program. (I have previously written of how similar perverse incentives plague the new CAEP accreditation standards for these institutions).

Here’s a test question for Arne Duncan, our Secretary of Education:
TRUE OR FALSE?
“A program’s ability to train future teachers who produce positive results in student learning [as measured by standardized testing] is a clear and important standard of teacher preparation program quality.” (from p. 63 in proposed regulations document)
Here’s a hint, provided by Benjamin Campbell of Richmond,Virginia on the federal register of comments. “Current research indicates that no more than 14% -- and often far less – of a student’s learning as measured by standard tests – the only standardized measure – can be attributed to the teacher.”

The bad news is that Arne Duncan, and a whole slew of politicians and policymakers in line behind him, think the correct answer to this question is TRUE. They actually believe harsh punitive consequences work and lead to improvement. They think closing schools and teacher education programs is a good idea. They don’t care if any of their plans are based on faulty data, junk science, or illogical statistics. They blithely ignore extant research, recommendations from experts, and, to put it bluntly, common sense. The question remains – what are we going to do about it?


As Captain Jean-Luc Picard would say, “Engage.”

Saturday, August 17, 2013

What’s wrong with evaluating teacher education programs in New York City?


Nothing, in theory. The NYCDOE has just released a report comparing the traditional teacher preparation programs at local institutions that provide over half of teachers for the public schools in the city. There isn’t tremendous variation when you look across the charts, and when you read the fine print and discover the caveats about the reliability of the data, as Valerie Strauss of the Washington Post did, you might just shrug off the whole thing as no big deal.

There are valid arguments to be made about the need for this sort of study. For starters, those considering teaching as a profession have a right to know if the program they choose will prepare them well and help them to find employment, among other things. It would also be good if there were ways to make the decision about which program to choose that went beyond comparing transit and billboard advertising, flashy brochures, tuition costs, and courses and credits needed for program completion. For example, students searching for a teacher preparation program might like to know some fairly important things such as the quality of the professors and adjuncts, the size of the classes, the presence of helpful advisors, the library resources, course scheduling and flexibility, the opportunity to choose electives or have options among requirements, the schools in which they will do fieldwork and student teaching, the teachers who will mentor them in those schools, just to name a few features. There’s no YELP for these things, and Rate My Professors isn’t much help either. There is always word of mouth. For example, I heard some young women on the subway commenting on the merits of the program at Relay. One was explaining to another, “I don’t even have to go to class. You pass these tests and you get an exemption from attending.” (Relay’s website refers to the possibility of “substitutions” for time in class). There were two national accreditation organizations, NCATE and TEAC, which have merged into CAEP under a new set of standards, but the transition will be murky and will take years, plus many in the field feel that everybody can get accredited and it’s not a particularly meaningful measure of quality.
 



There are the various rankings, starting with the most well-known, the U.S. News and World Report, Petersons, Forbes, even Standard & Poors. These pick a mix of variables and criteria and weigh them in different ways to line up the institutions from best to worst. Most recently, a controversial report prone to errors, missing data, and other validity issues from NCTQ & US News & World Report gave mostly scathing marks to all but four institutions preparing teachers. It’s evident the study was seriously flawed, and critics have pointed out that it is not so useful to just look at inputs such as syllabi and other material gathered from the internet (see Stanford Professor Linda Darling-Hammond’s commentary and Professor Aaron Pallas from Columbia University’s Teachers College humorous take here, or a consolidated debunking assembled by Professor Roxana Marachi here).


But in the era of accountability craziness, somebody decided that an important measure of teacher preparation programs was how well kids did on standardized tests in schools that hired graduates of the programs, and this has jumped to the head of the line in terms of output measures. There are plenty of problems with measuring teacher quality based on how kids do on standardized tests over time, known in jargon as VAM, or value-added measures, but the problems with connecting the dots to the teacher’s graduate program take the cake. The bias, invalidity issues, and even just the logistics make the problems exponentially greater. My brilliant advisor at the University of Michigan, Professor Virginia Richardson, noted back in 2007:

“So why argue against the use of student achievement scores as the outcome measure of  interest?  Probably because success on getting students to learn curriculum material—as measured by standardized tests is, by itself,  a simplistic, inadequate, misleading and perhaps immoral measure of teacher education quality. We must also look at the ways in which teachers conduct their classrooms, and at other elements that affect the success of teaching including such aspects such as willingness and effort by the learner, a social surround supportive of teaching and learning, and opportunity to teach and learn.”

What’s more, the tests that students are taking are in the process of changing and transitioning to be aligned with the Common Core State Standards. The latest results in New York City were so messed up as to be deemed invalid and worthless by many, yet they will become the basis for judgments about teacher performance, and now teacher education program quality. Even a veteran reporter like Beth Fertig had a hard time figuring out what to make of them, and testing expert Daniel Koretz wasn’t much help judging from this interview. Chancellor Merryl Tisch, in a radio interview with Brian Lehrer, gave vague answers to questions about how cut off scores were determined and about critiques from historian Diane Ravitch, who has recently called for the resignation of Commissioner John King on her blog.

What’s perhaps most worrisome is that this keen focus on student achievement as measured on tests and the misuse and distortion of that data is a growing trend nationally. In the New York Times article covering the report comparing the twelve colleges in New York City, the reporter ended with a quote from Secretary of Education Arne Duncan, who feels this study was a “major step forward.” I would add – into a bottomless pit.