Showing posts with label NCLB. Show all posts
Showing posts with label NCLB. Show all posts

Thursday, October 29, 2009

Measuring the Race to the Bottom

I've written extensively about the Race to the Bottom that is created by aspects of NCLB where states' performance is measured by how well each state meets its own targets. I've also pointed out that individual states participating in this race to the bottom are not particularly keen on having transparent ways to compare their standards with other states.

The National Center for Education Statistics (part of the Department of Education) has found a way to use data from the National Assessment of Educational Progress along side state accountability reports to actually examine and quantify any Race to the Bottom. In a new report, Mapping State Proficiency Standards Onto NAEP Scales: 2005-2007, they have looked at changes from 2005 to 2007 in state scores and how they compare with the national measure. The report looks at reading and math in the 4th and 8th grades.

A word about proficiency

The NEAP makes a distinction between a basic level and a proficient level of performance. For the NEAP proficient means competency over challenging subject matter and not merely grade-level performance. Most (all?) states also make a distinction on their state accountability tests. When talking about the TAKS test in Texas, the word proficient is often used to refer to the minimum passing requirement and the term commended is used to describe the higher level.

In Texas, parents will hear the word proficient to refer to the minimum standard of passing the TAKS. That is not how the word is used nationally. And it is not how I will use it here. I will try to avoid confusion where I can. But I suspect that the Texan use of the word proficiency is a form of grade inflation attempting to make families feel that, as in Lake Wobegon, in Texas all children are above average.

Comparison among states

The report compares state standards for the proficiency level (not the basic level). That is, this report, when it comes to Texas is looked at the level needed to score a commended TAKS result. It worked to determine what the NEAP cut-off would be for getting a commended result on the 4th and 8th grade math and reading TAKS. It did this for each state for which there was sufficient data. This allows us to compare the proficiency levels from state to state.

In 2007 data Texas falls below the national average in its commended levels for 4th and 8th grade math and reading. For Texas is fourth from the bottom in 8th grade reading, beating out only North Carolina, Georgia and Tennessee. (Note that DC, Nebraska and Utah weren't included in this measure due to insufficient data.) For 8th grade math, Texas is near the middle of the pack. For 4th grade reading and math, Texas falls near the top of the bottom third.

States with higher proficiency (commended) standards have few students meeting those standards. There should be no surprise there. This leads to the question of whether it matters at all where states set their proficiency standards. Remember that proficiency standards are higher than the basic standards which all students are expected to meet. It turns out that states that set their own higher proficiency standards appear to get better results on the national NAEP exams. Whether the setting of higher standards is the cause of those higher scores is unknown. It should be noted that this relationship is much less pronounced for 8th grade reading, where it is not statistically significant.

Comparison over time

The question we asked with respect to any race to the bottom is whether states are lowering their own standards over time. The rest of the report concerns comparing 2005 and 2007 data. Getting the comparisons is mathematically tricky and so is the statistical inferencing. The report discusses their techniques in great detail, which I have yet to carefully review.

For each of 4th and 8th grade math and reading, they did two kinds of comparisons. The first is simply looking at the NEAP scores corresponding to the commended cut-offs has changed from 2005 to 2007. In this, Texas had no real change in 4th or 8th grade reading or 4th grade math (there was a decline in NEAP points, but that was within the margin of error for the analysis). But for 8th grade math there was a statistically meaningful decline of 4.2 points on the NEAP scale.

The report also looked at change in state standards in another way. If a state had a large increase in the number of students reaching the commended (proficient) level from 2005 to 2007 but did not have such a large (or any) increase in numbers of students improving on the NEAP.

Using this measure Texas students showed significantly more improvement on the Texas tests than on the national tests in 4th grade reading, 4th grade math, and 8th grade math.

Are the state standards getting easier

The pattern of change describe for Texas can be seen in many states (while other states are going in other directions). But does this means that states are lowering their standards in a race to the bottom? It certainly could mean that, but I suspect that this is more a consequence of schools getting better at preparing students for the state tests.

Schools are teaching test taking skills that are geared to the state tests. They are providing hot breakfasts on test days, they are perfecting their ways of motivating students and families to perform well on these tests. And with the actual teaching of content, there may be an increase in teaching to the test. A great deal of these efforts to improve state test scores will not carry over to the NEAP tests. The state accountability tests are very high stakes tests for the schools, while the NEAP tests have little direct consequence for the students, teachers or schools.

So schools will be engaging in activities that improve state test performance but do little for NEAP tests. This way we can see the results reported without it meaning that states are formally lowering their standards. Of course, if I am right about this, it means that we should be even more skeptical of improvements in state test results. It doesn't reflect a real increase in learning, but instead improvements in taking the state tests.

Thursday, September 24, 2009

Promising noises from the Secretary of Education

Secretary of Education Arne Duncan was interviewed by the Christian Science Monitor and made some very promising remarks regarding NCLB in my opinion. There was nothing even approximating specifics, but I think that he hit on a key insight:

[Duncan] hopes to essentially turn the law on its head. The Bush administration’s legislation, he says, kept the goals loose but the steps tight. He hopes instead to see a law that keeps the goals tight but the steps loose.

Here Duncan is referring to the fact that NCLB very tightly monitors how each state meets its own (loose) standards. These can lead to what I and others have called a race to the bottom between states, particularly when states work to avoid comparison of their education standards.

Exactly how an overhaul of NCLB will tighten or provide some uniformity of the goals is not something I know. I can imagine a range of mechanisms each with their own advantages and problems.

Set a national curriculum
The problems with this are legion. I won't dwell on them other than to say there is little reason to believe that the federal government would do a better job at this than even the worst of our fifty states.
Provide interstate comparisons to parents
When parents get accountability information about their child's school and their child's test scores, simply have these compared to national norms. If state officials can no longer hide their state's performance from parents, that might be enough to get states to start racing to the top. A difficulty with this is that it may require even more testing of students using a nationally normed test. There may be technical ways to get comparable data that won't involve more testing, but it will take some thinking about. Another difficulty with this approach is that it the parental pressure it generates will be insufficient to do the job. Finally, we know that it is parents in the upper middle class who exert the most political pressure, but even in lagging states their children will probably be performing above the national norm.

Some combination of those and other things may be part of what gets proposed. I eagerly await the plan. As for loosening the controls on exactly how states meet the (tighter) goals I can't even begin to speculate. From the philosophical point of view, Duncan's remarks seem very promising and sensible. Although I have no idea of how to achieve this, I am looking forward to more specific announcement.

Thursday, September 17, 2009

Thinking about assessment

The education literature likes to make a distinction between assessment for learning and assessment of learning. The distinction is, in my view, a necessary insight, but the way that it is conceived is both too limiting and prone to confusion. In this rant I am going present a somewhat richer framework for discussing different types of assessment for different purposes.

Where I'm coming from

As I've mentioned before, I am training to be a high school math teacher, and I am enrolled in what I consider to be an outstanding program through Collin College. I must confess that when I signed up for the program, I, in my arrogance, did not think that I would learn much. I am pleased to report that I was dead wrong. I won't go into why I was wrong, but I will say that I go to bed thinking about the ideas that come up from class discussion and readings and I wake up thinking about them. I remain (very) critical of some of the argumentation and scholarship in the readings, but it is extremely helpful for me to read them. I'm gobbling them up and loving it.

I have been, and remain, highly critical of the kinds of testing and incentive systems that have been set up by NCLB even though I fully support the goal of keeping schools and districts accountable for how well they serve all students, particularly the ones who are at risk of being left behind. Please see my previous posts on the matter (and more to come). NCLB does appear to be reaching that stated goal but it distorts the educational system as a whole and hinders progress in other important areas. But this essay is about assessment (testing and similar things). Whether you are a critic or supporter of NCLB you will agree that it is has greatly intensified the amount and importance of (standardized) testing in schools.

The Educators' Complaint

The education literature makes a distinction between assessment of learning and assessment for learning. A similar distinction is also called summative assessment and formative assessment. I will not attempt to give a full definition of these here. I don't think that the definitions in the literature bear up under close inspection, and the fuller the definition the less enlightening it is. Instead here is the rough idea through examples. Assessment of includes things like the TAKS, end of term exams, and major examinations that determine a student's grade. Assessment for learning is the on-going assessment that teachers engage while teaching. These include asking questions of the class, seeing what sorts of questions students ask. These are considered for learning because they help the teacher adapt teaching to the particular student.

The problem with our increased emphasis on assessment of learning is that most of that assessment isn't pedagogically useful. Some even argue that it is harmful in and of itself beyond the misdirection of resources (although I have my doubts about that claim). NCLB is a reality (which really does appear to be meeting its narrow, but important, goals), but the concern among educators is that it leads to too much pedagogically useless assessment. I agree, but I think that we are talking about assessment in a far too limiting framework.

Distinguishing distinctions

When we look at assessment, and try to categorize it, I think that we need to be looking at two dimensions, instead of the one-dimensional approach in the of-for distinction. We need to ask

  1. What is the form of the assessment?
  2. What is the purpose of the assessment?

The current discussion seems to think that all standardized tests (form) serve only to assess what a student has learned and not to adjust teaching (purpose), while all of the less formal (form) assessments are only used to adjust teaching (purpose). Certainly there is a strong connection between form and function, but when looking at assessment it will be useful to look at these along these two not-quite-independent dimensions.

Three purposes

When it comes to considering the various purposes of assessment I think that it is helpful to consider three separate purposes, not just the two in the existing conceptualization.

  1. Adjusting: to help adjust teaching to the needs of the particular student
  2. Grading: to provide feedback to student and family, to assign grades and work as an incentive
  3. Accounting: to evaluate the teaching of the teacher, school, district.

Accounting is what we see in the testing that follows from NCLB. It is about rating and evaluating schools and districts (and within districts it will be used to evaluate teachers). It is the school administrators who have the most to gain or lose by these test results. And they are typically done at the end of the school year. Although students who fail the test will be intensively tutored so that they will pass a retake, these tests are not used to help students directly.

Grading is typically the assessments that a course grade is based upon. These are presented to parents and students. These become part of a student's record and are intended to indicate how much the student learned. Of course these will also feed back on how a particular student is taught. A teacher can learn from these that a student is not meeting expectations and so can look for ways to help the student. One characteristic of grading assessment is that it (almost) never goes beyond what has been taught in class.

Adjusting is used primarily to help determine how to teach a particular student. These can range from everyday queries while teaching to see if students are getting it or not. But at the other extreme these can be the kinds of evaluations that are used to determine whether a student should be in a gifted and talented program or in special education. Those typically involve highly formalized exams, but are used exclusively for determining how best to teach an individual student. Homework may be part of a student's grade (usually to get them to do it), but is used primarily as a frequent check of whether something needs to be retaught.

Any particular assessment can (and often) will serve multiple purposes. But when looking at any particular assessment it is useful to keep those three purposes in mind.

Form follows function except for when it doesn't

If you've been talking about the differences between similes and metaphors in class you may ask for examples to help with the learning that day (adjusting). But you may also ask for examples of each on an end of term examination (grading). So the same form can be used for different purposes in different contexts. I've praised the MAP testing that PISD does. But I honestly don't know what they use it for. I would hope that they use it to help differentiate teaching (adjusting), but it may be used primarily to track teacher performance (accounting). So here is a particular standardized test administered exactly the same way could be used for entirely different purposes.

Some forms of assessment really are single purpose. Some like the Texas TAKS tests can't be used for much other than accounting, and then only a limited type. The test is designed to distinguish between students who have acquired the basic knowledge expected for the grade level from those who have not. It doesn't do a very good job of discriminating between students at the high end or very low end. It is hard for me to imagine a set of exams that is more narrowly focused on one purpose.

With understanding come solutions

This understanding of purposes can bring real, practical, recommendations. The TAKS serves little direct pedagogical purpose other than accounting, we could save a great deal of time and money (that could then go to actually improving education) by sampling. Not every student needs to take the TAKS in every subject. Consider fifth grade TAKS requirements. Students take Reading, Math and Science. Not counting make-ups and such, that takes three full days for the students' to complete. But if the goal is to measure a schools' performance, then have one third of the students take Reading, one third Math, and one third Science. Students would be randomly assigned with neither student nor school staff knowing which student gets which test until test day. All of the tests can then be given on the same day.

I believe that the framework I've introduced above, first separating form from purpose and then distinguishing three separate purposes for assessment, allows for a more useful discussion of assessment than is common. At least it helps me think about these things more carefully, and I hope it does the same for any readers I might have.

Friday, September 11, 2009

Congratulations: I may be wrong

I have been ranting (particularly here and here) that Texas' implementation of NCLB is doing a disservice to above average students. I am perplexed, but delighted, to report evidence that I have been wrong. Apparently, Texas students have been making remarkable gains in passing Advanced Placement exams.

The TEA has reported strong gains in AP pass rates, and the gains among some minority groups are truly spectacular. Being the cynic that I am, I had first assumed that the results were a consequence of fewer students taking the exams. But, according to the report, these gains while the number of students taking the exams has increased. So this positive result does not (immediately) look like the result of statistical manipulation.

My skepticism remains, and there are a few things to check out. But at the moment we have some good news, and I will take it as such.

It will be interesting to learn how these gains were distributed throughout the State. Do they come from a few school districts, and are those districts doing something unusual? If anyone knows where I can get this data, please let me know.

Tuesday, September 8, 2009

No Child Gets Ahead - The evidence

In April I wrote a piece No Child Gets Ahead in which I argued that current implementations (and particularly in Texas) of the No Child Left Behind program is detrimental to the interests of the above average student. Let me also remind everyone that I consider the goals of NCLB laudable and important. Again, see that earlier rant for a defense of those goals.

Now there is increasing evidence that I am correct. Brighter students are not advancing at the rate one might normally expect of them. This was discussed in a New York Times opinion piece titled Smart Child Left Behind on August 28, 2009. The authors, Tom Loveless and Micheal Petrilli, refer at first to report by the Center for Educational Policy published in June 2009.

The Rosy CEP report

The CEP report asks the question in its title, Is the emphasis on proficiency shortchanging higher- and lower-achieving students? Their answer is no. But Loveless and Petrilli argue that the CEP's report is deeply flawed. After reading the report, I entirely agree that it is broken beyond repair. The most egregious error in that study is the exclusive use of state proficiency test scores. State proficiency tests are designed to measure skill at the grade proficiency level. They never test anything above grade level (which is where advanced students are). My anecdotal experience is that pretty much everyone in my daughter's gifted and talented program hit the ceiling (score 100%) of the state TAKS tests. State proficiency tests are not designed to measure learning beyond the grade level proficiency levels, and simply don't work to measure learning for the high level students. The CEP report pretty much spells out the flaw without realizing it

The main measure of student achievement for this study consists of data from the state tests in reading (or English language arts) and mathematics used for NCLB accountability. Although no large-scale test provides a complete picture of student achievement, we have analyzed state test results because these tests are given to nearly all students in a state, are intended to reflect each state’s academic content standards, and are designed to assess whether students have met their states’ expectations for performance at a particular grade level. [Emphasis mine.]

It appears that the CEP report measures success in a state by looking at state results in terms of the percentage of students (with each state) scoring proficient or above. This notion of counting the number of students who exceed a certain (minimal) standard as a way of seeing whether you are serving the higher performing students is entirely missing the point of the exercise. The question we are asking is Does the way we measure school success shortchange the top students?. The CEP's answer appears to be, Well if we count success according to the way NCLB measure it, then we have success..

Another astounding flaw in the CEP analysis is their use of states as their level of analysis. For them, a gain in a small state completely off sets a lose in a large state, even if it means a decline for millions of more students than there is a gain for. This is truly blushworthy error. even though they fully acknowledge it (on page 18). I could go on. These are not minor technical quibbles. These problems completely and utterly undermine the CEP conclusions.

Why worry

Before I go on to cite the evidence for my assertion that NCLB does shortchange the better students, let me spell out why I and so many others worry that it would do exactly that. I've outlined these reasons in my earlier rant and when combined with the actual level of these proficiency standards (see my rants, Race to the Bottom and No Comparison) there really is a concern. As I said before, people and systems do respond to incentive systems, so we should look very clearly at what we incentivize.

Two Scenarios

The NCLB incentive system rewards schools and districts for the number of students who pass (minimal) proficiency tests. The margin of passing or failing (how high above or below the passing cut-off) counts for nothing. Imagine a class with three students: Alice, Bob, and Charlie. And suppose that the proficiency level is considered met if a student scores a 70 on the crucial test. (Obviously I'm grossly simplifying the examples for the purposes of illustration.)

Now consider scenario 1: Alice scores a 78 and passes. Bob scores and 70 and passes, Charlie scores a 50 and fails. In this scenario, the class has two passes and one failure. That is what will be counted in determining the school's, district's, and (probably) teacher's rating.

Now consider scenario 2: Alice scores a 95 and passes, Bob scores a 69 and fails, Charlie scores a 62 and fails. This class has one pass and two failures. The school, district, and teacher will be marked down severely for this.

All of the incentives (and they are powerful incentives) of NCLB push for scenario 2 above scenario 1. But let's look which class is serving the students better. Both Alice and Charlie do much better in scenario 1 than they do in scenario 2. While Bob does slightly worse in scenario 1 then in 2.

Of course you may object that I could have set up an example where the class that did better on NCLB criteria would also be the one that we would all agree better served the students. But my example illustrates real choices that schools and teachers make every day.

Suppose that you are a teacher and you are confident that Alice will pass the exam with little extra effort from you. With more effort from you, she might learn a great deal, but she is already on a clear target to pass the exam. And suppose that Bob is a student who looks like he will pass the exam, but only if you put extra effort into preparing him. Finally, as a teacher, with all of your pre-tests and such, you determine that even with an extraordinary effort on your part, Charlie is unlikely to pass the exam. If you want to keep your job, and the school wants to keep property values high in its district, then you will focus your effort on Bob.

Response to Response to Intervention

In the excellent teacher training program that I am currently enrolled in, we have been studying the mechanisms by which we identify and help the struggling student, Bob, before he falls too far behind. It is a program (or framework) called Response to Intervention (RtI). It really looks like it should be effective at identifying students like Bob earlier and getting the teacher to devote more time to Bob's needs. But of course any additional time spent on Bob will be time taken away from Alice and Charlie (unless additional staff are provided or the school day is lengthened).

This is one thing that people always seem to forget. Anyone who says that we need to spend more time doing X (where X is "with struggling students", "in the library", "teaching math", "practicing bus evacuations", "taking tests", etc) needs to remember that that means spending less time doing something else. This applies to money as well as time. An additional dollar spent on X is a dollar taken away from something else. It's easy to say what we should spend more time or money on, but it's very hard to answer the question of where that time or money comes from.

How to find out

I've already explained that if we want to test whether the incentives set up by NCLB create a disservice to above average students, we can't measure that by counting how many states have an increase in the number of students reaching the proficient level in that state. So how do we check? First of all we will need to use measures that are (a) comparable across states, and (b) which accurately measure the skills of the above average student. Ideally, we would like to have (c) where the progress of individual students from year to year is measured.

Getting data that is comparable across states is difficult. NCLB allows each state to set its own minimum and proficient standards. Because both politicians and educators like to be able to boast about how well their students are doing, there is pressure to set these standards low. (There are also some good reasons to set them low.) As a consequence of this, there is an incentive to shy away from mechanisms that allow state standards to be compared with one another or have students from one state compared with those of another. (See my earlier rant, No Comparison.)

We also need achievement results that don't suffer from a ceiling effect. That is, it should assess the full range of student achievement including those students near the top. This can be difficult for a number of reasons. First of all, the state assessments for NCLB are completely unsuited for this; so any tests would need to be in addition to those required for NCLB. Secondly, most testing to see whether students have learned the material presented in class; thus they rarely can test students who are above grade level.

Fortunately, there have been an number of attempts to collect such data. In an earlier post, I discussed the Measure of Academic Progress produced by the Northwest Evaluation Association. This time, I will be looking at The Nation's Report Card: Writing 2007 produced by the National Center for Education Statistics (part of the US Department of Education). They developed a scale which should include most advanced students, and sampled school children from across the country. Details of their method can be found in the report. For our purposes, merely showing a chart on page 9 of their report should make the point.

nations report card writing 2007-page9.jpg

Before NCLB went into effect nation wide (2002) there was no growth in 8th grade writing skills at the lowest levels, while there were gains at the highest level. After NCLB went into effect, there were gains significant at the lowest levels and stagnation at the upper levels. Now I admit that I did troll through reports to find the most dramatic example. But for all grades studied and in all areas we find that NCLB has led wonderful gains at the lower levels. These are important and valued achievements. At the same time, it has lead to a flattening of growth at the higher levels.

In all fairness?

As I've said elsewhere, gains in one place often have costs elsewhere. If we have to have a trade off of improvements for the top students or improvements for the bottom students, maybe we are redressing a prior imbalance by focussing on the struggling student. I will address this issue in a later post. Here I will say that the situation before NCLB was destructive and unjust, with the below average abandoned. NCLB needs to be credited with fixing that. But the current situation, in which the above average student is ignored by the educational system, is little better. But whether you think that the current situation is right or wrong, I hope that everyone realizes that it does shortchange the above average students. In future posts, I will try to elaborate on how I think we can develop an accountability system that establishes incentives which serve all students.

Wednesday, June 24, 2009

No Comparison

Texas is refusing to join an effort to develop national standards for Math and English education. Texas, Alaska, Missouri, and South Carolina are the only states to decline. The stated reasons for refusing to participate and declining these Race to the Top Funds is cost and maintaining independence. The cost excuse doesn't hold water since a cost that is currently borne entirely within the state would be shared among many. The second reason, distaste for adopting any idea that wasn't developed in Texas, may well be sincere but is hardly helpful. My contention is that the real reason for refusing to participate is something else altogether: Texan politicians don't want a transparent comparison of our schools' achievements with those of other states.

What many people fail to recognize about the current No Child Left Behind program is that is measures how well schools and districts meet their own state standards. So states which set low proficiency standards will find that they perform better on NCLB measures than states that set a higher bar. And because each state develops its own testing, there is no easy way to see which states set the bar lower than others. This makes it possible for politicians in states to tout their achievements with the federal NCLB (making it seem to voters that this is a real national comparison) even as standards and results remain low. People in the state, wanting to believe that their state is holding their own, are eager to believe their politicians. This leads to a Lake Wobegon effect, where every state is above average.

I have argued earlier that there is nothing wrong with low minimum standards as long as they used as minimum standards instead of as targets. But the current system, in which each states sets its own target and then is judged on how well it meets that, just encourages a race to the bottom. What we have now obfuscates comparison among states, but we are going to break out of this race to the bottom, we need relatively easy and reliable ways for the public to compare education in the various states. So let's not let the State of Texas' pride and independence stand in the way of creating an education system that we can honestly be proud of.

Wednesday, April 1, 2009

In defense of low standards

I have in several places pointed out that what Texan's seem to think are high standards in education are typically very low when compared to the standards used by other states in the USA. And I will continue to do this for as long as I feel that people in Texas don't grasp how low the standards really are. But my argument here is that for some purposes, low standards are absolutely appropriate.

If we want to set some educational proficiency standard that we seriously expect all (or the overwhelming majority of) children to meet then we have to recognize that there are real differences in capabilities among individuals. That minimum standard should be well below what the average individual can achieve. The crucial fact about education is that one size does not fit all, but we do have to ensure that every child achieves some minimal proficiency. For this we need to allow minimum standards to be minimum.

Of course we should expect much more than the minimum from most children. A school in which every child meets a true minimum but little more is certainly failing to serve its students. Unfortunately the incentives in the current implementation of the No Child Left Behind program largely do direct schools to try to achieve the minimum for most students and provide little incentive to go beyond that. The temptation among many reformers is to raise the minimum standards. Unfortunately that will just have the consequence of leaving more children far behind, either through drop-outs or reclassification of children into exempt categories. Let's keep minimum standards as minimum standards.

If we want to establish incentives for schools to help all children, not just those near the minimum then we need to add something to the mix. We should measure schools by the year-on improvement of all children. If a high achieving student (who already meets the minimum standard) improves greatly, that should be to the school's credit. Likewise, a low achieving student (who doesn't meet the minimum standard) improves greatly than that should also be the the school's credit. Likewise, lack of improvement for either should count against the school. With incentives like these, then every child will be most improved irrespective of whether they also reach the minimum standard.

So the bottom line is that it is a good idea to keep the minimum standards minimum as long as those standards don't become the primary standards by which schools are judged.

Monday, February 9, 2009

No child gets ahead

Before I begin my rant about the No Child Left Behind program, let me state clearly four views that I enthusiastically share with the program.
  1. Every child should become at least minimally proficient in certain essential skills
  2. It is import to provide a way of comparing how well schools and districts achieve educational goals
  3. Incentives to schools and districts do affect school policies.
  4. Standard tests are the least subjective way of measuring students' progress
Looking at that list, you might think that I was also an enthusiastic supporter of NCLB. But you would be wrong.
What the above leads to is a system where incentives to schools, school districts, property owners and so on is the portion of students in a school who meet a minimum standard. The overwhelming portion of what goes into a school's rating is whether a sufficient portion of the children reach a minimum standard.
Now let's look at the consequence of this
  • Effort going to serve the educational needs of children who will comfortably exceed the minimum standard will be wasted. That is, improvement among students who would meet the minimum standard in the first place will not be reflected in school ratings.
  • Effort going to serve students who are unlikely to meet the minimum standard is also wasted in that improvements there are not reflected in the ratings.
  • Success of a school, district or state becomes measured in terms of minimal standards, thus curriculum will be geared toward the minimum standard.
As mentioned in my points of agreement above, incentives do work. Therefore, we need to check very carefully what kind of incentives we establishes. All of the negative consequences that I've listed are the result of the incentive system that NCLB creates.
How not to fix the system.
The system that we have in place provides us (at best) with measures of how well schools have students meet minimum standards. The worst thing that we could do is to treat that measure as anything other than what it is. It is not a measure of how well the school educates students on average; it is not a measure in any real sense of how "good" a school is. The only thing that this measure tells us is whether or not a school is failing to get most of its students to meet a minimal standard.
But because this is the only measure that we have for comparing schools and districts within a state, it gets used for purposes that it was never designed to serve. A test of how many students meet a minimal standard can only be used to identify failing schools. It is entirely useless at distinguishing decent schools from excellent schools. For example, in a successful district like Plano Independent School District, more students hit the ceiling (reach the maximum score possible on the test) than actually fail it [get source for this. It was buried in one of the NWEA documents]. The Texas Assessment of Knowledge and Skills (TAKS) completely fails to provide a measure that can be used to compare those who well exceed its minimum standard.
A natural reaction and proposed solution is to raise the tests' standards: Make the test harder. But that would actually make matters worse. If we raise the standards of the test to a degree where it would be meaningful for schools that are doing well within the state, then it no longer works as ensuring that all students reach a minimum standard. It will place the higher standard out of reach of some students. The fact of the matter is that not all students are alike. Not all students are college bound. And when we establish truly minimum standards we need to take that into account. If our "minimum" is no longer truly minimum and so become out of reach for some children, then those children will be left behind.
Of course we should set high goals for everyone so that we get the best from each, but we shouldn't insist that everyone reach high goals.
Possible fixes
I have several ideas for how to improve the system. I will sketch a few of them below, but will need to expand on them at some later date.
Make every score count
Instead of reporting just the number of students who merely pass the test, also report the average score. This way an improvement for any student (whether well below passing or well above passing) gets credited to the school. A simple average may not be the best number because of how outliers affect the results, but some statistic or set of statistics that makes every child's performance count will help motivate schools to help each and every child whether they are near the passing threshold or not.
But this can't be done with the test as it stands (at least in Texas). As I mentioned, the Texas test has a very low ceiling. In my district more students score 100% on each test than fail it. This means that the test is providing absolutely no useful measurement for those students. Also when scores reach a ceiling they have a perverse affect on averages. Developing a test which has the appropriate scoring system is difficult, but achievable. But I will leave that for another post as well.
Distinguish between pedagogically useful and useless tests
Tests like the TAKS provide little information to the teacher about how to help an individual child. The tests are not pedagogically useful. That's fine because that is not their intent. They are designed to help us compare schools and districts. Unfortunately a great deal of school, teacher, and student effort goes into pedagogically useless activity. There are several ideas of how to deal with this, all involve less individual testing. One would be to test less frequently, and another would be to test only a sample of the students in a school instead of testing the entire school population.
Simplify administration of tests by eliminating the conflict of interest
The rules and procedures that schools have to follow for administering the tests are beyond belief. Visit your local school and ask to see the printed guidelines. You won't have time to read them, and you may not even have time to count the pages. Just weigh them on a bathroom scale.
The reason for many of these rules is because there is a truly awful design decision in the administration of the tests. The people who have the most at stake (the teachers and the school officials) are the ones who are asked to administer the test. This is a massive conflict of interest. And if you believe, as I do, that incentive systems do affect how people behave, then you see that there is a terrible opportunity for test administrators helping students cheat on the tests.
Of course I believe that most teachers are honest, but I also lock my car even though I think that most people wouldn't steal it. If you ask any teacher about cheating on these tests they will of course tell you that it doesn't happen in their school, but they will also be familiar with some terrible abuses in other schools or districts. We have set up a systematic administrative conflict of interest and so add boatloads of Band-Aids to cover a wound that is wider than a church door and deeper than a well.
The entire administration of these non-pedagogical exams could be simplified if we had them administered by some third party.
I will try to expand on these thoughts and fill in many of the blanks in further posts. But at this point, I should just upload this one even though it fall far short of what I had hoped for it.