Today Judge McDonough ruled that Sheri Lederman had met the burden of proof for showing that her APPR VAM-based "ineffective" rating was "indisputably arbitrary and capricious." The recent actions of the New York State Regents to impose a four year moratorium on the use of VAM in teacher evaluations had led the state to seek a settlement with Lederman, but she held on for this important victory. The judge did rule that the second category of relief was moot by the actions of the Regents. What will be most useful moving forward is the language about inherent bias in the use of VAM. Judge McDonough wrote in his 15 page summary posted by Leonie Haimson that he found:
"...convincing and detailed evidence of VAM bias against teachers at both ends of the spectrum (e.g. those with high-performing students or those with low-performing students)....and most tellingly...a 'bell curve' that places teachers in four categories via pre-determined percentages regardless of whether the performance of students dramatically rose or dramatically fell from the previous year" (p. 11).
I wrote about the hearing and Judge McDonough's difficulties with bell curve logic back in August. It seems increasingly that the public is understanding the many problems with standardized, normed testing and the inappropriate ways it is being used. Evidence is mounting that national tests such as PARCC are created to produce high rates of failure and are not even aligned with the common core standards. See for example this recent account of a teacher revealing in detail inappropriate content and questions on the 4th grade PARCC posted by Teachers College professor Celia Oyler on her blog. Alan Singer also reported on his Huffington Post blog about high numbers of parents and students opting out of the state tests, and outrage about the content and difficulty level.
This is a day for celebrating, but the truth is, this testing nonsense is not going away anytime soon. Time to open our eyes and make our outrage known.
Showing posts with label VAM. Show all posts
Showing posts with label VAM. Show all posts
Tuesday, May 10, 2016
Thursday, August 13, 2015
It’s All About the Bell Curve: Sheri Lederman’s Day in Court
I traveled up to Albany this morning to hear the oral
arguments in the Lederman v. King case presented to Acting Supreme Court
Justice Roger McDonough by Bruce Lederman, and Colleen Galligan representing
the State Education Department. This is the first time in my life I have sat in
a courtroom proceeding. I don’t even watch Law and Order. Let’s just say I was
most definitely not in my element. But I’m a pretty good observer of human
behavior, a decent note-taker, and I had personal reasons for caring deeply
about the outcome of this case, above and beyond all the reasons we all should
care about a case that may have far-reaching implications for the misguided
reforms of Race to the Top (see full disclosure below). What I witnessed was a
masterful take down of the we-need-objectivity rhetoric that is plaguing
education. So I should begin by saying that I am hopeful, because it seems
someone with the power to make a difference gets it. Judge McDonough gets that
it’s all about the bell curve, and the bell curve is biased and subjective.
In case you need a refresher on how test scoring works these
days (and who doesn’t) I suggest you start with the excellent fact sheets from
Fair Test, first on norm-referenced tests, or NRTs, and then on criterion-referenced tests, or CRTs, and tests used to measure performance against state standards. In particular
note the following important points:
“NRTs are designed to
sort and rank students 'on the curve,' not to see if they met a
standard or criterion. Therefore, NRTs should not be used to assess whether
students have met standards. However, in some states or districts a NRT is used
to measure student learning in relation to standards. Specific cut-off scores
on the NRT are then chosen (usually by a committee) to separate levels of
achievement on the standards. In some cases, a CRT is made using technical
procedures developed for NRTs, causing the CRT to sort students in ways that
are inappropriate for standards-based decisions.”
As you may notice,
we’ve come a long way from getting a 91 out of 100 on a test and knowing that
was an A-. Testing today is obtuse and confusing by design. In New York State, we boil it down to a ranking
from one to four. That’s right, there’s even jargon for “ones and twos” that is
particularly heinous when you learn that politicians have interests in making
more than 50% of students fall in those “failing” categories. Today the state
released the test score results for students in grades 3-8 and their so-called
“proficiency” is reported as below 40% achieving the passing levels. By design
the public is meant to read this as miserable failure.
The political
narrative of public education failure extends next to the teachers, who must
demonstrate student learning based on these faulty tests, even if they don’t
teach the subjects tested, and even if they teach students who face hurdles and
hardships that have a tremendous impact on their ability to do well on the
tests. In Sheri’s case, her rating plunged from 14 out of 20 points to 1 out of
20 points on student growth measures. Yet her students perform exceedingly well
on the exams; once you are a “four” you can’t go up to a “four plus” because
you’ve hit the ceiling. In fact, one wrong answer could unreasonably mark you
as a “three” and you would never know. Similarly, the teacher receives a
student growth score that is also based on a comparison to other teachers. When
it emerged in the hearing today that the model, also known as VAM, or
value-added, pre-determined that 7% of the teachers would be rated
“ineffective” Judge McDonough caught on to the injustice that lies at the heart
of the bell curve logic: where you rank in the ratings is SUBJECTIVE.
In his affidavits, Professor Aaron Pallas of Teachers College brilliantly explains the many
flaws with this misuse of student test scores to evaluate and rank teachers’
effectiveness. Predetermining a set percentage of ineffective teachers
regardless of their actual “effectiveness” and their students’ achievements was
the first major flaw. The second is that the model is not grounded in
scientific definitions of teacher quality or effectiveness, as there are many
factors beyond a teacher’s control that contribute to student performance on
standardized tests and other measures of their knowledge and skills. Third, the
model is not transparent on what “needs to be done to achieve effective or
highly effective ratings” which is a requirement of the law. The model also
violates the law’s definition of student growth as “change in student
achievement for an individual student between two or more points in time.”
Judge McDonough seemed to have picked up on this idea, and asked if a better
model would test the student at the start and end of a given academic year.
Pallas gives a far more nuanced explanation of the need for a different model
of testing to measure growth over time, but suffice it to say, the model that
produced Sheri’s absurd score is not measuring student growth as defined by the
law. Pearson, the corporate entity behind the testing enterprise, even noted,
“It is inappropriate to compare scale scores across grades as they neither
measure the same content, nor are they on the same scale.” Yet that is what the growth model does.
The lame explanation
from Colleen Galligan was that the model may not be perfect but the state tries
to compare each student to similar students. The goal, she offered, is to find
outliers in the teaching pool who consistently have a pattern of
ineffectiveness, to either give them additional training or fire them. At this
point Judge McDonough offered her a chance to explain the dramatic drop in
Sheri’s score. “On its face it must mean students bombed the test (speaking as
one who has bombed tests)” and this produced laughter in the courtroom. For who
hasn’t bombed at least one test in their life? Who has not experienced that
dread and fear of being labeled a failure? Then Judge McDonough asked
rhetorically, “Did they learn nothing?” The only answer she could come up with,
was that in this case Dr. Lederman’s students, although admittedly performing
well compared to other students, did worse than 98% of students across the
state in growth. At this point it was pretty clear to everyone present that
this made absolutely no sense whatsoever.
Full disclosure:
Sheri Lederman is my high school classmate and she is a
highly regarded elementary teacher in the Great Neck Public Schools, which we
both attended in our childhoods. She got her doctorate at Hofstra University,
where my mother is a professor emerita, and where I know many of the faculty as
personal friends. They confirm the high regard I have for Sheri’s intelligence
and insights into education. I think she is absolutely heroic to be pursuing a
lawsuit, with the expert guidance of her lawyer husband, Bruce Lederman,
against the New York State Department of Education, to expose the irrational
and illegal practices of evaluating teacher performance using “arbitrary and
capricious” student growth models based on flawed science. I have previously written in my blog about Sheri’s hope that her lawsuit would prove to be a
“tipping point” in halting the use of these erroneous student growth models. A bit of background on the case from last October can be
found here.
On June 1st, the New York State Supreme Court
ruled that Sheri’s case could go forward despite the State Education
Department’s claim that her lawsuit was baseless since Sheri’s overall
evaluation was “effective” despite the “ineffective” label on the student
growth portion, worth 20% of the total.
Today’s news was covered so far here, here and here. The local CBS station covered it here and WNYT here.
Wednesday, April 1, 2015
Cause for hope. That's right, hope!
A decade ago, Marilyn
Cochran-Smith, then president of AERA, gave us a portrait of teacher education
at a major crossroads – for better or for worse – and invited us to see where
these divergent paths might lead us. At the time, in that hotel ballroom, her bleak portrayal of
a future hostile to the ideals and values so many of us held, brought me to
tears as I thought of my graduate students who were making so many sacrifices
and working so hard to become exemplary urban educators in schools others had
written off as “failing” or “ghetto” and doomed.
Why then today, when the draconian education
reform ideas of Governor Cuomo have succeeded in a budget vote last night in
Albany, do I feel there is cause for hope? Because last night I attended a
lecture at Teachers College by the brilliant scholar and public intellectual
David C. Berliner. Known especially for his exhaustive defense of public education,
The Manufactured Crisis, a 1995 book co-authored with Bruce J. Biddle, Berliner
is a master of reasoned argument and robust evidence, all presented with
clarity and that wow factor that makes you wonder how anyone could possibly
disagree with him. His latest book, 50 Myths & Lies that Threaten America's Public Schools: The Real Crisis in Education, written with Gene Glass, is a must read. He began by characterizing current efforts to evaluate
teachers and teacher education using standardized test scores of students as
“ridiculous” and proceeded to enumerate a dozen slam dunk arguments against
these misguided reforms. I was reminded of the spell Harry Potter learned from
Professor Lupin to cast on a boggart, Riddikulus, transforming the scary into
the humorous. Below is a photo of the final slide summarizing his reasons.
![]() |
| Berliner, D. C. (2015) "Evaluating Teachers and Teacher Education Using Student Test Data: A Misunderstanding" |
The
Teachers College Sachs Lecture Series has set out to explore teacher education
– its future, worth, and change (both reactive and transformative) – by
examining curriculum, current practices, beliefs about knowledge, and needed
research. Prior to Berliner, Marilyn Cochran Smith spoke of education reform, a
“hot and huge topic” she defined as “a set of practices and policies, many of
which were set into motion with NCLB, although with deeper roots, and continued
or accelerated by RTTT. These were intended to fix America’s broken education
system keeping with a neoliberal, market-based approach and a heavy focus on
accountability.” She described the current state of teacher education as at
best, uneven, and at worst, uninspired, ineffective, and out of touch.
Many of the pessimistic ideas outlined in her AERA address
ten years ago are now in motion. We are getting closer to a national database
to track teacher education programs and their impact on student test scores,
rewarding those effective in test results and enabling them to become, in her
words, “lucrative national franchises.” Just look at the expansion of Relay. Politicians
and policymakers continue to believe that creating competition and ranking
programs, rewarding winners and sanctioning or closing down losers, is going to
lead to improvement.
What
concerns me is that current ideas about teacher quality are terribly
ill-defined, in large part because much of the research is based on statistics
of standardized test results, which are horrible proxies for anything
meaningful, and as Berliner pointed out, they measure next to nothing about
teacher effects. Right now we are stuck, trapped in the bad idea that all that
matters is successful teaching defined as making test scores go up. We have
lost all regard for whether good teaching matters, whether the ends to means
are ethical and justifiable. Driven by data Data DATA, it seems there is no
trust in human judgment, or testimonials of stakeholders who can speak
passionately to the difference teachers have made in their lives. They only
want numbers.
Take for
example the latest incarnation of the bad ideas in teacher education reform,
the Deans for Impact. Based in Austin, Texas with a million dollar start up
grant, a group of deans from various colleges of education across the country
have resolved to be “data driven, outcomes focused, transparent and
accountable” and to use “empirically tested” features in their programs that
improve student learning. On their website they boast, “We want to inject some
of the values of start-up culture into higher education.” Their first order of
business was to write a support statement in early March for the new accreditation
organization, CAEP, stating “Deans for Impact stands ready to bring all hands
on deck to help CAEP succeed.” Mercedes Schneider has already done the
necessary investigation to connect the reform dots and money behind this
venture.
As Jorge Cabrera wrote recently, we are witnessing “a form a social engineering under the guise of ‘urgency’
and ‘reform’” and it comes as no surprise that these deans want to speed up the
phase-in of new federal regulations by two years. A rush to implementation will
create exactly the sort of chaos and havoc that allows them to ramp up the rhetoric
of failure. Their litany of complaints is all too familiar: “teacher prep” is
awful, there’s too much theory and not enough practice, it’s too easy to become
a teacher. Backed by the simplistic critique of Arthur Levine, these ideas have
paved the way for erroneous experiments of throwing beginners to the wolves
with little more than a few weeks of boot camp preparation.
But the
hopeful side has some promise and I believe that we have reached a tipping
point. The misuse of value-added measures, or VAMs, in evaluating teachers and
teacher education programs is poised for some harsh pushback. The AERA
publication, Educational Researcher, has dedicated its latest issue to the VAM controversy. In a succinct and lucid
editorial by my former professor at the University of Michigan, Stephen
Raudenbush, he deplores the distorted use of VAMs. He cautions, “The hard
question is how to integrate the new research on teachers with other important
strands of research in order to inform rather than distort practical judgment.”
He goes on to pose the question, “Does the answer to a precisely focused
research question, by itself, have
implications for practical action?” Aside from consistent and reliable
evidence, he argues the need for a powerful theory of action to synthesize all
of the evidence.
Get ready
for some intense work ahead of us. They will continue to put lipstick on the
pig with slick ads, propaganda and celebrity endorsements, and splashy rallies,
but money and political power can only go so far. Parents, teachers,
professors, and students of all ages must work together on their common educational
goals to restore sanity for the good of our universities, schools, and
communities.
Friday, January 2, 2015
Standardized Testing: The Final Frontier
Tests seem so reasonable at first -- teachers teach, students learn, and
demonstrate mastery by passing a test. But as Daniel Koretz says at the start
of his 2008 book, Measuring Up: What Educational Testing Really Tells Us, “Achievement testing is a very complex
enterprise, and as a result, test scores are widely misunderstood and misused.”
Now that is what I call an understatement. Furthermore, despite Common Core
claims that better standards and tests mean fewer reasons for concern about
their misuse, as Vito Perrone of Harvard University pointed out, “Most items on
these various standardized tests remain well within the longstanding technology
of testing, primarily to support the mechanical scoring procedures. They still
seem to be limited instruments with too much influence” (1999, p. 152).
The testing “enterprise” is poised for a warp drive
record-breaker of misuse insanity. In a nutshell, here’s how they plan to
connect the dots.
A tiny fraction of what a student knows and can do is
hypothetically captured, with some modicum of so-called scientific accuracy, by
converting the number of correct answers out of the total number of questions
on a standardized test to a raw score. Keep in mind that this single raw test
score is still prone to error in its intent to measure what the student knows
as the student may have made random choices, guessing correctly (or not), or
may simply have had other contextual reasons for the performance including
illness, distraction, nerves, etc. The test is also imperfect by design and is
likely biased in some ways.
Now that raw score goes through some psychometric process to
either be normed to a scale comparing it to other test scores, and/or it is
ranked somewhere between unacceptable and excellent based on someone’s judgment
of what students should know and be able to do. This is where all hell breaks
loose as that converted score gets used.
How might it get used? For one, to tell the students and the
parents or guardians how “well” they did which can involve labeling the
converted score with a percentile rank, a grade-level equivalent, or just a
descriptive meaning such as “meets standard.” However, it will likely be used
in what is called a “high stakes” way to assign students to special education,
to hold them back a year, or to track them into homogenous groups.
The most pernicious use is to group the scores to make
claims about the quality of individual teachers. From there, it’s easy to see
how tempting it is to make a claim about the quality of a school, and then a
whole district. While we’re at it, let’s compare counties, states, regions,
countries.
The cold hard truth, in Koretz’s words, is this:
Scores on a single
test are now routinely used as if they were a comprehensive summary of what
students know or what schools produce (p. 44-45).
He goes on later to add:
Simply attributing differences in
scores to school quality or, similarly, simply assuming that scores themselves
are sufficient to reveal educational effectiveness, is unrealistic. And more
generally, simple explanations of performance differences are usually naïve.
All of this is established science (p. 142).
Things get really tricky when hierarchical linear modeling
kicks in to provide a “value-added” way to compare actual scores to a
prediction and to use the difference to rate teachers’ effectiveness. Ignoring
warnings from experts, these value-added models or VAMS, have been misused by
policymakers to weigh heavily in the annual evaluation of teachers. Carol
Burris, an outspoken principal who opposed this misuse of standardized test
scores, recently wrote of a teacher’s lawsuit filed in New York State by my
friend, Sheri Lederman, who hopes her case can become “a tipping point” in
bringing this damaging unreliable practice to a grinding halt.
That may be wishful thinking because now the dots are being
connected to the colleges and universities that educate teachers. They too are
to be evaluated and ranked based on the performance of their candidates for
teacher certification on standardized tests, which can be more than four in
some cases. New federal regulations currently open for public comment until February 2nd would
require these institutions of higher education to also track their teacher
graduates, and collect their annual evaluation ratings including the VAMS
measure, in order to be considered eligible for the TEACH grant program. (I
have previously written of how similar perverse incentives plague the new CAEP
accreditation standards for these institutions).
Here’s a test question for Arne Duncan, our Secretary of
Education:
TRUE OR FALSE?
“A program’s ability to train future teachers who produce
positive results in student learning [as measured by standardized testing] is a
clear and important standard of teacher preparation program quality.” (from p. 63 in proposed regulations document)
Here’s a hint, provided by Benjamin Campbell of Richmond,Virginia on the federal register of comments. “Current research indicates that
no more than 14% -- and often far less – of a student’s learning as measured by
standard tests – the only standardized measure – can be attributed to the
teacher.”
The bad news is that Arne Duncan, and a whole slew of
politicians and policymakers in line behind him, think the correct answer to
this question is TRUE. They actually believe harsh punitive consequences work
and lead to improvement. They think closing schools and teacher education
programs is a good idea. They don’t care if any of their plans are based on
faulty data, junk science, or illogical statistics. They blithely ignore extant
research, recommendations from experts, and, to put it bluntly, common sense.
The question remains – what are we going to do about it?
As Captain Jean-Luc Picard would say, “Engage.”
Thursday, September 5, 2013
Teacher Preparation Standards Add New Outcomes
The
new standards for accrediting teacher preparation programs have just been adopted by
the Board of Directors of the Council for the Accreditation of Educator
Preparation (CAEP) to
ensure that candidates completing programs are “classroom-ready” and can demonstrate they have a positive impact on student learning. This
latter quality seems a reasonable goal for a new teacher, and there is careful
language about using multiple measures of valid and reliable data that is
spelled out to include ways of measuring growth over time. The first three
standards all address what happens during the program, but standard four is a new addition, an attempt to measure candidates’
effectiveness and success in their careers, and their satisfaction with their
professional preparation, once they have graduated and have been teaching in
full time jobs. It’s going to mean that institutions will have to do more than
keep a list of alumni, they will have to continue to gather complex data and
create survey instruments, following graduates even when they move out of
state.
So although it may seem
reasonable to say that institutions preparing teachers should be held
accountable for the outcomes of that preparation, there are simply too many
factors outside the control of the institution for that to be a reasonable
requirement, at least in the way it has been designed in the CAEP standards.
The other big problem, which underlies so much of what is wrong with
educational policy, is the misuse of standardized testing to measure things it
was not designed to measure. Serious concerns have been raised about the use of
value-added measures (VAMs) seeking to capture student growth over time, and
using those scores as a heavy consideration in rating teachers and granting
tenure (see this 2011 letter from top experts to the New York State Regents for example).
Now these teacher preparation standards are asking even more of the VAMs, and
it is quite an unreasonable and dangerous leap to say they are valid measures
of the quality of teacher preparation programs.
We are engrossed in a policy world in teacher education that doesn’t think through the unintended consequences of making changes to standards and reporting requirements of institutions. Extant research on the statistical invalidity of using value-added measures of student learning and growth to evaluate teachers, schools, and programs is ignored. There is little or no discussion of the moral implications of these policy decisions, and there is no way to know whether the ends to means are achieved in defensible ways or through cheating and gaming the system. As Diane Ravitch has said, “Test-based accountability encourages a slew of negative behaviors.” It also creates burdensome and costly administrative requirements to collect and analyze data that do little to promote meaningful reforms to programs and improve how and what teacher educators teach their students. These standards create requirements that will suck time and money away from teacher preparation institutions, and in teaching, time and money are precious resources that should be carefully protected and closely monitored.
Let’s just make an analogy for a moment. Let’s say we
imagine two doctors who finish medical school, pass the board exams, and get
full time jobs in the emergency rooms of hospitals. One works in a large city ER,
while the other is in a suburban ER. The medical school will be judged on the
health records of the patients of these two doctors, as well as their
satisfaction with their education. The city ER doctor has twice the patients of
the other doctor, horrible working conditions, constant stress from treating
gunshot and stab wound victims related to gang activity, and feels betrayed by
her medical education that did not prepare her for this reality. She eventually
quits. The other doctor enjoys an efficiently run hospital with well-trained
colleagues, and the personal rewards of helping his patients get immediate care
and treatment, and saving lives. He gives high marks on a survey measuring his
satisfaction with his medical education. The medical school is penalized in the
case of the ER doctor, and rewarded in the case of the other doctor. As cases
like this accumulate, the program shifts from encouraging and preparing doctors
to work in settings like the first one to creating more opportunities for
doctors to go to jobs like the second one. How might such a trend negatively
impact the health of patients?
Now let’s imagine a case of two teachers. They are from the
same institution, and are enrolled in masters programs that lead to dual
certification in students with disabilities. The first teacher is getting
certified in early childhood, the second teacher in secondary mathematics. They
both graduate with honors, pass the state exams with high scores, and find full
time jobs. The first is employed in a large city district in a first grade
inclusion classroom in a high poverty school. Her students are not old enough
for state tests, so the first problem is what will be the valid measures of
this teacher’s impact on her students’ learning? How can we control for
absenteeism, student turnover, neglect or abuse in the home, the effects of
homelessness, hunger, and other common conditions of poverty? What of the
effect of working conditions on this teacher, such as a large class size,
stress, lack of resources or support, staff and administrative turnover? How
will we know if she uses harsh punishments to “manage” poor behavior? Suppose
she has poor evaluations from her principal because her students often misbehave
and are not reading and writing well enough. When asked to rate her teacher
preparation program in a survey, she gives low marks because she feels she was
not prepared for the realities of her current position. After two years, she
quits.
The other teacher gets a job in a suburban district teaching
special education classes to ninth graders who are way behind and hate math.
Their scores from previous state tests are very low. This teacher uses computer
programs to provide remediation, and rewards students frequently with prizes,
candy, free time and pizza parties to improve their attitudes. Although they
love the teacher and their test scores do go up a little, they really haven’t
improved their knowledge or skills enough to handle high school level algebra.
The teacher does have positive evaluations from the principal because of the
perception that students are well behaved, participate, and make progress on
the computer programs. This teacher gets tenure and gives high marks to the
institution that prepared him.
The teacher preparation program is penalized for not meeting
the requirements of standard four in the case of the first teacher, and
rewarded for meeting them in the case of the second. As cases like this
accumulate, the institution shifts focus from encouraging and preparing
teachers like the first one to preparing more teachers like the second one. Over
time, the focus in teacher preparation is on effectiveness, defined as performance on standardized tests by the
teacher’s students, as the singular measure of teacher quality, and concerns
for how teachers attend to moral dimensions of learning, or to issues of
diversity, diminish. How might such a trend negatively impact learning for all
students?
We are engrossed in a policy world in teacher education that doesn’t think through the unintended consequences of making changes to standards and reporting requirements of institutions. Extant research on the statistical invalidity of using value-added measures of student learning and growth to evaluate teachers, schools, and programs is ignored. There is little or no discussion of the moral implications of these policy decisions, and there is no way to know whether the ends to means are achieved in defensible ways or through cheating and gaming the system. As Diane Ravitch has said, “Test-based accountability encourages a slew of negative behaviors.” It also creates burdensome and costly administrative requirements to collect and analyze data that do little to promote meaningful reforms to programs and improve how and what teacher educators teach their students. These standards create requirements that will suck time and money away from teacher preparation institutions, and in teaching, time and money are precious resources that should be carefully protected and closely monitored.
Subscribe to:
Posts (Atom)

