Showing posts with label Sheri Lederman. Show all posts
Showing posts with label Sheri Lederman. Show all posts

Tuesday, May 10, 2016

Good News About the Lederman Lawsuit on APPR

Today Judge McDonough ruled that Sheri Lederman had met the burden of proof for showing that her APPR VAM-based "ineffective" rating was "indisputably arbitrary and capricious." The recent actions of the New York State Regents to impose a four year moratorium on the use of VAM in teacher evaluations had led the state to seek a settlement with Lederman, but she held on for this important victory. The judge did rule that the second category of relief was moot by the actions of the Regents. What will be most useful moving forward is the language about inherent bias in the use of VAM. Judge McDonough wrote in his 15 page summary posted by Leonie Haimson that he found:

"...convincing and detailed evidence of VAM bias against teachers at both ends of the spectrum (e.g. those with high-performing students or those with low-performing students)....and most tellingly...a 'bell curve' that places teachers in four categories via pre-determined percentages regardless of whether the performance of students dramatically rose or dramatically fell from the previous year" (p. 11). 

I wrote about the hearing and Judge McDonough's difficulties with bell curve logic back in August. It seems increasingly that the public is understanding the many problems with standardized, normed testing and the inappropriate ways it is being used. Evidence is mounting that national tests such as PARCC are created to produce high rates of failure and are not even aligned with the common core standards. See for example this recent account of a teacher revealing in detail inappropriate content and questions on the 4th grade PARCC posted by Teachers College professor Celia Oyler on her blog.  Alan Singer also reported on his Huffington Post blog about high numbers of parents and students opting out  of the state tests, and outrage about the content and difficulty level.

This is a day for celebrating, but the truth is, this testing nonsense is not going away anytime soon. Time to open our eyes and make our outrage known. 

Thursday, August 13, 2015

It’s All About the Bell Curve: Sheri Lederman’s Day in Court

I traveled up to Albany this morning to hear the oral arguments in the Lederman v. King case presented to Acting Supreme Court Justice Roger McDonough by Bruce Lederman, and Colleen Galligan representing the State Education Department. This is the first time in my life I have sat in a courtroom proceeding. I don’t even watch Law and Order. Let’s just say I was most definitely not in my element. But I’m a pretty good observer of human behavior, a decent note-taker, and I had personal reasons for caring deeply about the outcome of this case, above and beyond all the reasons we all should care about a case that may have far-reaching implications for the misguided reforms of Race to the Top (see full disclosure below). What I witnessed was a masterful take down of the we-need-objectivity rhetoric that is plaguing education. So I should begin by saying that I am hopeful, because it seems someone with the power to make a difference gets it. Judge McDonough gets that it’s all about the bell curve, and the bell curve is biased and subjective.

In case you need a refresher on how test scoring works these days (and who doesn’t) I suggest you start with the excellent fact sheets from Fair Test, first on norm-referenced tests, or NRTs, and then on criterion-referenced tests, or CRTs, and tests used to measure performance against state standards. In particular note the following important points:

“NRTs are designed to sort and rank students 'on the curve,' not to see if they met a standard or criterion. Therefore, NRTs should not be used to assess whether students have met standards. However, in some states or districts a NRT is used to measure student learning in relation to standards. Specific cut-off scores on the NRT are then chosen (usually by a committee) to separate levels of achievement on the standards. In some cases, a CRT is made using technical procedures developed for NRTs, causing the CRT to sort students in ways that are inappropriate for standards-based decisions.”

As you may notice, we’ve come a long way from getting a 91 out of 100 on a test and knowing that was an A-. Testing today is obtuse and confusing by design. In New York State, we boil it down to a ranking from one to four. That’s right, there’s even jargon for “ones and twos” that is particularly heinous when you learn that politicians have interests in making more than 50% of students fall in those “failing” categories. Today the state released the test score results for students in grades 3-8 and their so-called “proficiency” is reported as below 40% achieving the passing levels. By design the public is meant to read this as miserable failure.

The political narrative of public education failure extends next to the teachers, who must demonstrate student learning based on these faulty tests, even if they don’t teach the subjects tested, and even if they teach students who face hurdles and hardships that have a tremendous impact on their ability to do well on the tests. In Sheri’s case, her rating plunged from 14 out of 20 points to 1 out of 20 points on student growth measures. Yet her students perform exceedingly well on the exams; once you are a “four” you can’t go up to a “four plus” because you’ve hit the ceiling. In fact, one wrong answer could unreasonably mark you as a “three” and you would never know. Similarly, the teacher receives a student growth score that is also based on a comparison to other teachers. When it emerged in the hearing today that the model, also known as VAM, or value-added, pre-determined that 7% of the teachers would be rated “ineffective” Judge McDonough caught on to the injustice that lies at the heart of the bell curve logic: where you rank in the ratings is SUBJECTIVE.

In his affidavits, Professor Aaron Pallas of Teachers College brilliantly explains the many flaws with this misuse of student test scores to evaluate and rank teachers’ effectiveness. Predetermining a set percentage of ineffective teachers regardless of their actual “effectiveness” and their students’ achievements was the first major flaw. The second is that the model is not grounded in scientific definitions of teacher quality or effectiveness, as there are many factors beyond a teacher’s control that contribute to student performance on standardized tests and other measures of their knowledge and skills. Third, the model is not transparent on what “needs to be done to achieve effective or highly effective ratings” which is a requirement of the law. The model also violates the law’s definition of student growth as “change in student achievement for an individual student between two or more points in time.” Judge McDonough seemed to have picked up on this idea, and asked if a better model would test the student at the start and end of a given academic year. Pallas gives a far more nuanced explanation of the need for a different model of testing to measure growth over time, but suffice it to say, the model that produced Sheri’s absurd score is not measuring student growth as defined by the law. Pearson, the corporate entity behind the testing enterprise, even noted, “It is inappropriate to compare scale scores across grades as they neither measure the same content, nor are they on the same scale.” Yet that is what the growth model does.

The lame explanation from Colleen Galligan was that the model may not be perfect but the state tries to compare each student to similar students. The goal, she offered, is to find outliers in the teaching pool who consistently have a pattern of ineffectiveness, to either give them additional training or fire them. At this point Judge McDonough offered her a chance to explain the dramatic drop in Sheri’s score. “On its face it must mean students bombed the test (speaking as one who has bombed tests)” and this produced laughter in the courtroom. For who hasn’t bombed at least one test in their life? Who has not experienced that dread and fear of being labeled a failure? Then Judge McDonough asked rhetorically, “Did they learn nothing?” The only answer she could come up with, was that in this case Dr. Lederman’s students, although admittedly performing well compared to other students, did worse than 98% of students across the state in growth. At this point it was pretty clear to everyone present that this made absolutely no sense whatsoever.
 
Sheri Lederman speaking to reporters outside the courtroom in Albany
Full disclosure:
Sheri Lederman is my high school classmate and she is a highly regarded elementary teacher in the Great Neck Public Schools, which we both attended in our childhoods. She got her doctorate at Hofstra University, where my mother is a professor emerita, and where I know many of the faculty as personal friends. They confirm the high regard I have for Sheri’s intelligence and insights into education. I think she is absolutely heroic to be pursuing a lawsuit, with the expert guidance of her lawyer husband, Bruce Lederman, against the New York State Department of Education, to expose the irrational and illegal practices of evaluating teacher performance using “arbitrary and capricious” student growth models based on flawed science. I have previously written in my blog about Sheri’s hope that her lawsuit would prove to be a “tipping point” in halting the use of these erroneous student growth models. A bit of background on the case from last October can be found here.

On June 1st, the New York State Supreme Court ruled that Sheri’s case could go forward despite the State Education Department’s claim that her lawsuit was baseless since Sheri’s overall evaluation was “effective” despite the “ineffective” label on the student growth portion, worth 20% of the total.

Today’s news was covered so far here, here and here. The local CBS station covered it here and WNYT here.




Friday, January 2, 2015

Standardized Testing: The Final Frontier

Tests seem so reasonable at first --  teachers teach, students learn, and demonstrate mastery by passing a test. But as Daniel Koretz says at the start of his 2008 book, Measuring Up: What Educational Testing Really Tells Us, “Achievement testing is a very complex enterprise, and as a result, test scores are widely misunderstood and misused.” Now that is what I call an understatement. Furthermore, despite Common Core claims that better standards and tests mean fewer reasons for concern about their misuse, as Vito Perrone of Harvard University pointed out, “Most items on these various standardized tests remain well within the longstanding technology of testing, primarily to support the mechanical scoring procedures. They still seem to be limited instruments with too much influence” (1999, p. 152).

The testing “enterprise” is poised for a warp drive record-breaker of misuse insanity. In a nutshell, here’s how they plan to connect the dots.

A tiny fraction of what a student knows and can do is hypothetically captured, with some modicum of so-called scientific accuracy, by converting the number of correct answers out of the total number of questions on a standardized test to a raw score. Keep in mind that this single raw test score is still prone to error in its intent to measure what the student knows as the student may have made random choices, guessing correctly (or not), or may simply have had other contextual reasons for the performance including illness, distraction, nerves, etc. The test is also imperfect by design and is likely biased in some ways.

Now that raw score goes through some psychometric process to either be normed to a scale comparing it to other test scores, and/or it is ranked somewhere between unacceptable and excellent based on someone’s judgment of what students should know and be able to do. This is where all hell breaks loose as that converted score gets used.

How might it get used? For one, to tell the students and the parents or guardians how “well” they did which can involve labeling the converted score with a percentile rank, a grade-level equivalent, or just a descriptive meaning such as “meets standard.” However, it will likely be used in what is called a “high stakes” way to assign students to special education, to hold them back a year, or to track them into homogenous groups.

The most pernicious use is to group the scores to make claims about the quality of individual teachers. From there, it’s easy to see how tempting it is to make a claim about the quality of a school, and then a whole district. While we’re at it, let’s compare counties, states, regions, countries.

The cold hard truth, in Koretz’s words, is this:
Scores on a single test are now routinely used as if they were a comprehensive summary of what students know or what schools produce (p. 44-45).
He goes on later to add:
Simply attributing differences in scores to school quality or, similarly, simply assuming that scores themselves are sufficient to reveal educational effectiveness, is unrealistic. And more generally, simple explanations of performance differences are usually naïve. All of this is established science (p. 142).

Things get really tricky when hierarchical linear modeling kicks in to provide a “value-added” way to compare actual scores to a prediction and to use the difference to rate teachers’ effectiveness. Ignoring warnings from experts, these value-added models or VAMS, have been misused by policymakers to weigh heavily in the annual evaluation of teachers. Carol Burris, an outspoken principal who opposed this misuse of standardized test scores, recently wrote of a teacher’s lawsuit filed in New York State by my friend, Sheri Lederman, who hopes her case can become “a tipping point” in bringing this damaging unreliable practice to a grinding halt.

That may be wishful thinking because now the dots are being connected to the colleges and universities that educate teachers. They too are to be evaluated and ranked based on the performance of their candidates for teacher certification on standardized tests, which can be more than four in some cases. New federal regulations currently open for public comment until February 2nd would require these institutions of higher education to also track their teacher graduates, and collect their annual evaluation ratings including the VAMS measure, in order to be considered eligible for the TEACH grant program. (I have previously written of how similar perverse incentives plague the new CAEP accreditation standards for these institutions).

Here’s a test question for Arne Duncan, our Secretary of Education:
TRUE OR FALSE?
“A program’s ability to train future teachers who produce positive results in student learning [as measured by standardized testing] is a clear and important standard of teacher preparation program quality.” (from p. 63 in proposed regulations document)
Here’s a hint, provided by Benjamin Campbell of Richmond,Virginia on the federal register of comments. “Current research indicates that no more than 14% -- and often far less – of a student’s learning as measured by standard tests – the only standardized measure – can be attributed to the teacher.”

The bad news is that Arne Duncan, and a whole slew of politicians and policymakers in line behind him, think the correct answer to this question is TRUE. They actually believe harsh punitive consequences work and lead to improvement. They think closing schools and teacher education programs is a good idea. They don’t care if any of their plans are based on faulty data, junk science, or illogical statistics. They blithely ignore extant research, recommendations from experts, and, to put it bluntly, common sense. The question remains – what are we going to do about it?


As Captain Jean-Luc Picard would say, “Engage.”