Showing posts with label TAKS. Show all posts
Showing posts with label TAKS. Show all posts

Friday, November 27, 2009

In Praise of the TAKS

A few days ago I had an eye-opening experience. I took the 2009 Exit level Math TAKS and have come to the conclusion that it is a far better designed test than I'd anticipated. Below I will explain both my initial pessimism and what impressed me about the test.

Before that, just some questions to get out of the way. I took the TAKS as part of a course assignment in my teacher training program. I missed one (out of 60) questions on an easy question due to a silly oversight. I didn't take the test under normal test taking conditions. On the one hand, I was free to make myself a fresh cup of tea every now and then; on the other hand, I had to put up with the dogs barking at the gardeners. I recommend that anyone – like me – who gripes about Texas standards and teaching to the test should try taking these exit level tests. Past tests are available from the Texas Education Agency for all levels and subjects for which the TAKS is administered.

Why I was pessimistic

I had not expected the TAKS exit level test to be as interesting as it turned out to be. There were several reasons for my pessimism. First was experience with the 3rd, 4th, and 5th grade math TAKS that my daughter took. They didn't really seem to involve any problem solving or mathematical reasoning; instead they were about the ability to apply memorized techniques to clear instances. In retrospect this was probably because educators take Piaget too seriously and (incorrectly) believe that grade school kids are incapable of formal reasoning. Whatever the cause, there is a large difference in approach between the grade school TAKS and the exit exam.

The second reason for my pessimism was based on my impression that in high school math education very little time is given to developing mathematical reasoning skills, and most of the time is on specific techniques to solve specific kinds of problems. Real understanding and creative problem solving rarely seemed to be emphasized. I had (incorrectly) attributed this to teaching to the test. I had thought what I had disliked about the curriculum was a consequence of the test.

Liking the test

Many problems on the test had multiple ways of getting at the solution. One method would be mindlessly applying the right set of procedures, plugging away at it (typically using a calculator), and eventually coming to an answer. But these problems were also set up as little mathematical puzzles. There was often a key insight which could lead one quickly, easily, and without a calculator to the right answer. Grasping the key to these problems not only saved time, effort, and tedium; but it also was less error prone. Many of the incorrect answer options were exactly the kinds of things one might arrive at for making a minor error (as I did in the question I got wrong). The more steps involved in computing an answer, the more opportunity there is to slip up on one of those common errors. If one had a good grasp of the meaning of the various mathematical concepts then the key was usually available.

Other questions were explicitly about concepts. These questions were not merely testing knowledge of technical vocabulary, but did require an understanding of the concepts to answer correctly. I don't think I have the skill to come up with questions of that nature, but I can recognize them when I see them.

As a minor anecdote that nicely illustrates how wrong I was about this test there was a question on the test that I had previously claimed would not be on a TAKS test. During my teacher training, I've taught some sample lessons to my fellow teachers-in-training. In one of them, I had students develop a model for an instances of n(n-1)/2 growth. I said at the time that this was an activity that would help them think mathematically but would not be on the TAKS. It was a dig at the TAKS and a completely unfounded one. Imagine my surprise when I hit question 6 on last year's test and discovered it to be exactly the kind of problem I said would not be on the TAKS.

I'm hoping that I will find time over the next few days to try exit exams in English, Science and Social Studies as well. Although high school teachers are specialists, we should – at a minimum – understand what is being expected of our students in all areas. (Now all I have to do is find someone who will grade the ELA writing sample if I do take that portion of that test.)

Decisions and Revisions

I was wrong in my expectations of the TAKS. Alternatively, I may have been correct in my expectations and wrong about my current evaluation of the TAKS. In either case I've been far off the mark at least once. While I certainly don't like being wrong, I actually enjoy the experience of discovering that I've been wrong. It is eye-opening in the best sense. I see things that I previously did not see, and I am forced to reevaluate both the reasoning that led up to the incorrect view and the consequences of that view.

But the big lesson in discovering that I've been spectacularly wrong about something is the heightened awareness that other positions that I currently hold firmly may also be wrong. Caveat lector.

Thursday, October 29, 2009

Measuring the Race to the Bottom

I've written extensively about the Race to the Bottom that is created by aspects of NCLB where states' performance is measured by how well each state meets its own targets. I've also pointed out that individual states participating in this race to the bottom are not particularly keen on having transparent ways to compare their standards with other states.

The National Center for Education Statistics (part of the Department of Education) has found a way to use data from the National Assessment of Educational Progress along side state accountability reports to actually examine and quantify any Race to the Bottom. In a new report, Mapping State Proficiency Standards Onto NAEP Scales: 2005-2007, they have looked at changes from 2005 to 2007 in state scores and how they compare with the national measure. The report looks at reading and math in the 4th and 8th grades.

A word about proficiency

The NEAP makes a distinction between a basic level and a proficient level of performance. For the NEAP proficient means competency over challenging subject matter and not merely grade-level performance. Most (all?) states also make a distinction on their state accountability tests. When talking about the TAKS test in Texas, the word proficient is often used to refer to the minimum passing requirement and the term commended is used to describe the higher level.

In Texas, parents will hear the word proficient to refer to the minimum standard of passing the TAKS. That is not how the word is used nationally. And it is not how I will use it here. I will try to avoid confusion where I can. But I suspect that the Texan use of the word proficiency is a form of grade inflation attempting to make families feel that, as in Lake Wobegon, in Texas all children are above average.

Comparison among states

The report compares state standards for the proficiency level (not the basic level). That is, this report, when it comes to Texas is looked at the level needed to score a commended TAKS result. It worked to determine what the NEAP cut-off would be for getting a commended result on the 4th and 8th grade math and reading TAKS. It did this for each state for which there was sufficient data. This allows us to compare the proficiency levels from state to state.

In 2007 data Texas falls below the national average in its commended levels for 4th and 8th grade math and reading. For Texas is fourth from the bottom in 8th grade reading, beating out only North Carolina, Georgia and Tennessee. (Note that DC, Nebraska and Utah weren't included in this measure due to insufficient data.) For 8th grade math, Texas is near the middle of the pack. For 4th grade reading and math, Texas falls near the top of the bottom third.

States with higher proficiency (commended) standards have few students meeting those standards. There should be no surprise there. This leads to the question of whether it matters at all where states set their proficiency standards. Remember that proficiency standards are higher than the basic standards which all students are expected to meet. It turns out that states that set their own higher proficiency standards appear to get better results on the national NAEP exams. Whether the setting of higher standards is the cause of those higher scores is unknown. It should be noted that this relationship is much less pronounced for 8th grade reading, where it is not statistically significant.

Comparison over time

The question we asked with respect to any race to the bottom is whether states are lowering their own standards over time. The rest of the report concerns comparing 2005 and 2007 data. Getting the comparisons is mathematically tricky and so is the statistical inferencing. The report discusses their techniques in great detail, which I have yet to carefully review.

For each of 4th and 8th grade math and reading, they did two kinds of comparisons. The first is simply looking at the NEAP scores corresponding to the commended cut-offs has changed from 2005 to 2007. In this, Texas had no real change in 4th or 8th grade reading or 4th grade math (there was a decline in NEAP points, but that was within the margin of error for the analysis). But for 8th grade math there was a statistically meaningful decline of 4.2 points on the NEAP scale.

The report also looked at change in state standards in another way. If a state had a large increase in the number of students reaching the commended (proficient) level from 2005 to 2007 but did not have such a large (or any) increase in numbers of students improving on the NEAP.

Using this measure Texas students showed significantly more improvement on the Texas tests than on the national tests in 4th grade reading, 4th grade math, and 8th grade math.

Are the state standards getting easier

The pattern of change describe for Texas can be seen in many states (while other states are going in other directions). But does this means that states are lowering their standards in a race to the bottom? It certainly could mean that, but I suspect that this is more a consequence of schools getting better at preparing students for the state tests.

Schools are teaching test taking skills that are geared to the state tests. They are providing hot breakfasts on test days, they are perfecting their ways of motivating students and families to perform well on these tests. And with the actual teaching of content, there may be an increase in teaching to the test. A great deal of these efforts to improve state test scores will not carry over to the NEAP tests. The state accountability tests are very high stakes tests for the schools, while the NEAP tests have little direct consequence for the students, teachers or schools.

So schools will be engaging in activities that improve state test performance but do little for NEAP tests. This way we can see the results reported without it meaning that states are formally lowering their standards. Of course, if I am right about this, it means that we should be even more skeptical of improvements in state test results. It doesn't reflect a real increase in learning, but instead improvements in taking the state tests.

Thursday, September 17, 2009

Thinking about assessment

The education literature likes to make a distinction between assessment for learning and assessment of learning. The distinction is, in my view, a necessary insight, but the way that it is conceived is both too limiting and prone to confusion. In this rant I am going present a somewhat richer framework for discussing different types of assessment for different purposes.

Where I'm coming from

As I've mentioned before, I am training to be a high school math teacher, and I am enrolled in what I consider to be an outstanding program through Collin College. I must confess that when I signed up for the program, I, in my arrogance, did not think that I would learn much. I am pleased to report that I was dead wrong. I won't go into why I was wrong, but I will say that I go to bed thinking about the ideas that come up from class discussion and readings and I wake up thinking about them. I remain (very) critical of some of the argumentation and scholarship in the readings, but it is extremely helpful for me to read them. I'm gobbling them up and loving it.

I have been, and remain, highly critical of the kinds of testing and incentive systems that have been set up by NCLB even though I fully support the goal of keeping schools and districts accountable for how well they serve all students, particularly the ones who are at risk of being left behind. Please see my previous posts on the matter (and more to come). NCLB does appear to be reaching that stated goal but it distorts the educational system as a whole and hinders progress in other important areas. But this essay is about assessment (testing and similar things). Whether you are a critic or supporter of NCLB you will agree that it is has greatly intensified the amount and importance of (standardized) testing in schools.

The Educators' Complaint

The education literature makes a distinction between assessment of learning and assessment for learning. A similar distinction is also called summative assessment and formative assessment. I will not attempt to give a full definition of these here. I don't think that the definitions in the literature bear up under close inspection, and the fuller the definition the less enlightening it is. Instead here is the rough idea through examples. Assessment of includes things like the TAKS, end of term exams, and major examinations that determine a student's grade. Assessment for learning is the on-going assessment that teachers engage while teaching. These include asking questions of the class, seeing what sorts of questions students ask. These are considered for learning because they help the teacher adapt teaching to the particular student.

The problem with our increased emphasis on assessment of learning is that most of that assessment isn't pedagogically useful. Some even argue that it is harmful in and of itself beyond the misdirection of resources (although I have my doubts about that claim). NCLB is a reality (which really does appear to be meeting its narrow, but important, goals), but the concern among educators is that it leads to too much pedagogically useless assessment. I agree, but I think that we are talking about assessment in a far too limiting framework.

Distinguishing distinctions

When we look at assessment, and try to categorize it, I think that we need to be looking at two dimensions, instead of the one-dimensional approach in the of-for distinction. We need to ask

  1. What is the form of the assessment?
  2. What is the purpose of the assessment?

The current discussion seems to think that all standardized tests (form) serve only to assess what a student has learned and not to adjust teaching (purpose), while all of the less formal (form) assessments are only used to adjust teaching (purpose). Certainly there is a strong connection between form and function, but when looking at assessment it will be useful to look at these along these two not-quite-independent dimensions.

Three purposes

When it comes to considering the various purposes of assessment I think that it is helpful to consider three separate purposes, not just the two in the existing conceptualization.

  1. Adjusting: to help adjust teaching to the needs of the particular student
  2. Grading: to provide feedback to student and family, to assign grades and work as an incentive
  3. Accounting: to evaluate the teaching of the teacher, school, district.

Accounting is what we see in the testing that follows from NCLB. It is about rating and evaluating schools and districts (and within districts it will be used to evaluate teachers). It is the school administrators who have the most to gain or lose by these test results. And they are typically done at the end of the school year. Although students who fail the test will be intensively tutored so that they will pass a retake, these tests are not used to help students directly.

Grading is typically the assessments that a course grade is based upon. These are presented to parents and students. These become part of a student's record and are intended to indicate how much the student learned. Of course these will also feed back on how a particular student is taught. A teacher can learn from these that a student is not meeting expectations and so can look for ways to help the student. One characteristic of grading assessment is that it (almost) never goes beyond what has been taught in class.

Adjusting is used primarily to help determine how to teach a particular student. These can range from everyday queries while teaching to see if students are getting it or not. But at the other extreme these can be the kinds of evaluations that are used to determine whether a student should be in a gifted and talented program or in special education. Those typically involve highly formalized exams, but are used exclusively for determining how best to teach an individual student. Homework may be part of a student's grade (usually to get them to do it), but is used primarily as a frequent check of whether something needs to be retaught.

Any particular assessment can (and often) will serve multiple purposes. But when looking at any particular assessment it is useful to keep those three purposes in mind.

Form follows function except for when it doesn't

If you've been talking about the differences between similes and metaphors in class you may ask for examples to help with the learning that day (adjusting). But you may also ask for examples of each on an end of term examination (grading). So the same form can be used for different purposes in different contexts. I've praised the MAP testing that PISD does. But I honestly don't know what they use it for. I would hope that they use it to help differentiate teaching (adjusting), but it may be used primarily to track teacher performance (accounting). So here is a particular standardized test administered exactly the same way could be used for entirely different purposes.

Some forms of assessment really are single purpose. Some like the Texas TAKS tests can't be used for much other than accounting, and then only a limited type. The test is designed to distinguish between students who have acquired the basic knowledge expected for the grade level from those who have not. It doesn't do a very good job of discriminating between students at the high end or very low end. It is hard for me to imagine a set of exams that is more narrowly focused on one purpose.

With understanding come solutions

This understanding of purposes can bring real, practical, recommendations. The TAKS serves little direct pedagogical purpose other than accounting, we could save a great deal of time and money (that could then go to actually improving education) by sampling. Not every student needs to take the TAKS in every subject. Consider fifth grade TAKS requirements. Students take Reading, Math and Science. Not counting make-ups and such, that takes three full days for the students' to complete. But if the goal is to measure a schools' performance, then have one third of the students take Reading, one third Math, and one third Science. Students would be randomly assigned with neither student nor school staff knowing which student gets which test until test day. All of the tests can then be given on the same day.

I believe that the framework I've introduced above, first separating form from purpose and then distinguishing three separate purposes for assessment, allows for a more useful discussion of assessment than is common. At least it helps me think about these things more carefully, and I hope it does the same for any readers I might have.

Tuesday, September 8, 2009

No Child Gets Ahead - The evidence

In April I wrote a piece No Child Gets Ahead in which I argued that current implementations (and particularly in Texas) of the No Child Left Behind program is detrimental to the interests of the above average student. Let me also remind everyone that I consider the goals of NCLB laudable and important. Again, see that earlier rant for a defense of those goals.

Now there is increasing evidence that I am correct. Brighter students are not advancing at the rate one might normally expect of them. This was discussed in a New York Times opinion piece titled Smart Child Left Behind on August 28, 2009. The authors, Tom Loveless and Micheal Petrilli, refer at first to report by the Center for Educational Policy published in June 2009.

The Rosy CEP report

The CEP report asks the question in its title, Is the emphasis on proficiency shortchanging higher- and lower-achieving students? Their answer is no. But Loveless and Petrilli argue that the CEP's report is deeply flawed. After reading the report, I entirely agree that it is broken beyond repair. The most egregious error in that study is the exclusive use of state proficiency test scores. State proficiency tests are designed to measure skill at the grade proficiency level. They never test anything above grade level (which is where advanced students are). My anecdotal experience is that pretty much everyone in my daughter's gifted and talented program hit the ceiling (score 100%) of the state TAKS tests. State proficiency tests are not designed to measure learning beyond the grade level proficiency levels, and simply don't work to measure learning for the high level students. The CEP report pretty much spells out the flaw without realizing it

The main measure of student achievement for this study consists of data from the state tests in reading (or English language arts) and mathematics used for NCLB accountability. Although no large-scale test provides a complete picture of student achievement, we have analyzed state test results because these tests are given to nearly all students in a state, are intended to reflect each state’s academic content standards, and are designed to assess whether students have met their states’ expectations for performance at a particular grade level. [Emphasis mine.]

It appears that the CEP report measures success in a state by looking at state results in terms of the percentage of students (with each state) scoring proficient or above. This notion of counting the number of students who exceed a certain (minimal) standard as a way of seeing whether you are serving the higher performing students is entirely missing the point of the exercise. The question we are asking is Does the way we measure school success shortchange the top students?. The CEP's answer appears to be, Well if we count success according to the way NCLB measure it, then we have success..

Another astounding flaw in the CEP analysis is their use of states as their level of analysis. For them, a gain in a small state completely off sets a lose in a large state, even if it means a decline for millions of more students than there is a gain for. This is truly blushworthy error. even though they fully acknowledge it (on page 18). I could go on. These are not minor technical quibbles. These problems completely and utterly undermine the CEP conclusions.

Why worry

Before I go on to cite the evidence for my assertion that NCLB does shortchange the better students, let me spell out why I and so many others worry that it would do exactly that. I've outlined these reasons in my earlier rant and when combined with the actual level of these proficiency standards (see my rants, Race to the Bottom and No Comparison) there really is a concern. As I said before, people and systems do respond to incentive systems, so we should look very clearly at what we incentivize.

Two Scenarios

The NCLB incentive system rewards schools and districts for the number of students who pass (minimal) proficiency tests. The margin of passing or failing (how high above or below the passing cut-off) counts for nothing. Imagine a class with three students: Alice, Bob, and Charlie. And suppose that the proficiency level is considered met if a student scores a 70 on the crucial test. (Obviously I'm grossly simplifying the examples for the purposes of illustration.)

Now consider scenario 1: Alice scores a 78 and passes. Bob scores and 70 and passes, Charlie scores a 50 and fails. In this scenario, the class has two passes and one failure. That is what will be counted in determining the school's, district's, and (probably) teacher's rating.

Now consider scenario 2: Alice scores a 95 and passes, Bob scores a 69 and fails, Charlie scores a 62 and fails. This class has one pass and two failures. The school, district, and teacher will be marked down severely for this.

All of the incentives (and they are powerful incentives) of NCLB push for scenario 2 above scenario 1. But let's look which class is serving the students better. Both Alice and Charlie do much better in scenario 1 than they do in scenario 2. While Bob does slightly worse in scenario 1 then in 2.

Of course you may object that I could have set up an example where the class that did better on NCLB criteria would also be the one that we would all agree better served the students. But my example illustrates real choices that schools and teachers make every day.

Suppose that you are a teacher and you are confident that Alice will pass the exam with little extra effort from you. With more effort from you, she might learn a great deal, but she is already on a clear target to pass the exam. And suppose that Bob is a student who looks like he will pass the exam, but only if you put extra effort into preparing him. Finally, as a teacher, with all of your pre-tests and such, you determine that even with an extraordinary effort on your part, Charlie is unlikely to pass the exam. If you want to keep your job, and the school wants to keep property values high in its district, then you will focus your effort on Bob.

Response to Response to Intervention

In the excellent teacher training program that I am currently enrolled in, we have been studying the mechanisms by which we identify and help the struggling student, Bob, before he falls too far behind. It is a program (or framework) called Response to Intervention (RtI). It really looks like it should be effective at identifying students like Bob earlier and getting the teacher to devote more time to Bob's needs. But of course any additional time spent on Bob will be time taken away from Alice and Charlie (unless additional staff are provided or the school day is lengthened).

This is one thing that people always seem to forget. Anyone who says that we need to spend more time doing X (where X is "with struggling students", "in the library", "teaching math", "practicing bus evacuations", "taking tests", etc) needs to remember that that means spending less time doing something else. This applies to money as well as time. An additional dollar spent on X is a dollar taken away from something else. It's easy to say what we should spend more time or money on, but it's very hard to answer the question of where that time or money comes from.

How to find out

I've already explained that if we want to test whether the incentives set up by NCLB create a disservice to above average students, we can't measure that by counting how many states have an increase in the number of students reaching the proficient level in that state. So how do we check? First of all we will need to use measures that are (a) comparable across states, and (b) which accurately measure the skills of the above average student. Ideally, we would like to have (c) where the progress of individual students from year to year is measured.

Getting data that is comparable across states is difficult. NCLB allows each state to set its own minimum and proficient standards. Because both politicians and educators like to be able to boast about how well their students are doing, there is pressure to set these standards low. (There are also some good reasons to set them low.) As a consequence of this, there is an incentive to shy away from mechanisms that allow state standards to be compared with one another or have students from one state compared with those of another. (See my earlier rant, No Comparison.)

We also need achievement results that don't suffer from a ceiling effect. That is, it should assess the full range of student achievement including those students near the top. This can be difficult for a number of reasons. First of all, the state assessments for NCLB are completely unsuited for this; so any tests would need to be in addition to those required for NCLB. Secondly, most testing to see whether students have learned the material presented in class; thus they rarely can test students who are above grade level.

Fortunately, there have been an number of attempts to collect such data. In an earlier post, I discussed the Measure of Academic Progress produced by the Northwest Evaluation Association. This time, I will be looking at The Nation's Report Card: Writing 2007 produced by the National Center for Education Statistics (part of the US Department of Education). They developed a scale which should include most advanced students, and sampled school children from across the country. Details of their method can be found in the report. For our purposes, merely showing a chart on page 9 of their report should make the point.

nations report card writing 2007-page9.jpg

Before NCLB went into effect nation wide (2002) there was no growth in 8th grade writing skills at the lowest levels, while there were gains at the highest level. After NCLB went into effect, there were gains significant at the lowest levels and stagnation at the upper levels. Now I admit that I did troll through reports to find the most dramatic example. But for all grades studied and in all areas we find that NCLB has led wonderful gains at the lower levels. These are important and valued achievements. At the same time, it has lead to a flattening of growth at the higher levels.

In all fairness?

As I've said elsewhere, gains in one place often have costs elsewhere. If we have to have a trade off of improvements for the top students or improvements for the bottom students, maybe we are redressing a prior imbalance by focussing on the struggling student. I will address this issue in a later post. Here I will say that the situation before NCLB was destructive and unjust, with the below average abandoned. NCLB needs to be credited with fixing that. But the current situation, in which the above average student is ignored by the educational system, is little better. But whether you think that the current situation is right or wrong, I hope that everyone realizes that it does shortchange the above average students. In future posts, I will try to elaborate on how I think we can develop an accountability system that establishes incentives which serve all students.

Wednesday, April 1, 2009

In defense of low standards

I have in several places pointed out that what Texan's seem to think are high standards in education are typically very low when compared to the standards used by other states in the USA. And I will continue to do this for as long as I feel that people in Texas don't grasp how low the standards really are. But my argument here is that for some purposes, low standards are absolutely appropriate.

If we want to set some educational proficiency standard that we seriously expect all (or the overwhelming majority of) children to meet then we have to recognize that there are real differences in capabilities among individuals. That minimum standard should be well below what the average individual can achieve. The crucial fact about education is that one size does not fit all, but we do have to ensure that every child achieves some minimal proficiency. For this we need to allow minimum standards to be minimum.

Of course we should expect much more than the minimum from most children. A school in which every child meets a true minimum but little more is certainly failing to serve its students. Unfortunately the incentives in the current implementation of the No Child Left Behind program largely do direct schools to try to achieve the minimum for most students and provide little incentive to go beyond that. The temptation among many reformers is to raise the minimum standards. Unfortunately that will just have the consequence of leaving more children far behind, either through drop-outs or reclassification of children into exempt categories. Let's keep minimum standards as minimum standards.

If we want to establish incentives for schools to help all children, not just those near the minimum then we need to add something to the mix. We should measure schools by the year-on improvement of all children. If a high achieving student (who already meets the minimum standard) improves greatly, that should be to the school's credit. Likewise, a low achieving student (who doesn't meet the minimum standard) improves greatly than that should also be the the school's credit. Likewise, lack of improvement for either should count against the school. With incentives like these, then every child will be most improved irrespective of whether they also reach the minimum standard.

So the bottom line is that it is a good idea to keep the minimum standards minimum as long as those standards don't become the primary standards by which schools are judged.

Saturday, March 14, 2009

Stopping the Race to the Bottom

In his March 10, 2009 speech on education, President Obama called on states to stop the "race to the bottom" with respect to educational proficiency standards. I will be proposing here a simple, inexpensive measure which will help stop that race to the bottom.

Race to the bottom

The race to the bottom is a simple consequence of the fact that politicians want parents to be happy about their children's achievements and these same politicians directly or indirectly control the proficiency standards. Setting standards low allows parents to be truthfully told that their children are exceeding standards. Furthermore, the national No Child Left Behind program specifically penalizes schools and districts which fail to meet their state standards, so again states are given a financial incentives to keep standards low.

I've argued elsewhere that there is a place for minimum standards, but those minimum standards should not be the reference point by which we measure our children's academic growth and achievements. So in this post I will make one small proposal that a district like PISD can do to meet the President's challenge and help families and the community focus on achievement beyond the minimum standard.

Some MAP background

In the elementary grades, the Plano Independent School District (Texas) uses the excellent MAP testing system developed by the Northwest Evaluation Association (NWEA). This test has many virtues which I will try to discuss in more detail other posts. But one virtue is that it is scored on an "equal interval scale" which means that a 5 point gain one year is comparable to a 5 point gain in another year. This makes the system ideal for measuring progress and changes in progress over time. Plano helpfully provides these charts (which they call "Learning Growth Charts" to parents during parent-teacher conferences on through the parent information system website. Here is a sample The orange line is the student's scores for the tests taken at different times. The band below it is what Plano ISD calls a "proficiency range." The MAP test can be given in the Fall, Winter and/or Spring during each school year, and the time marks along the bottom of the graph are the grade and season testing dates Most parents will see their children's lines well above the proficiency range and be pleased with the school and the district.

The problem

What parents aren't told very clearly is that the proficiency range is really just the range of scores near the TAKS passing estimates, and so in Texas, it is a very low range indeed. (More on just how low that is later). The whole system is rigged to conceal just how low the proficiency range really is, and to make (most) parents happy by exaggerating how well their children are doing.

But by making a (legitimate) minimum standard the reference point by which our kids' achievements and progress are measured we lower our goals and become satisfied with things that we really shouldn't be satisfied with.

A solution

One very simple solution to this very specific problem is to include the national (not state) norms for student scores on the same chart. This would allow parents to see how their child's achievements compare nationally and also show how the state proficiency level compares to the national average performance.

So there I propose a simple approach that could be implemented cheaply and easily for districts that already provide scores for tests with national data. Other districts, when the report students' performance against their own states proficiency measures can find ways to make clear where those standards stand against other state's standards and against national performance data.

Despite the simplicity and affordability of this proposal, I anticipate that people will find ways to resist this proposal. I intend to ask all of the candidates running for the PISD school board to comment on this. It will be interesting to see responses.

For comparison

The NWEA publishes national norm data (PDF) for the MAP, but I haven't been able to find information on the variance. They've also looked at scores on the MAP and seen which students passed the TAKS, thus being able to estimate just how hard or easy the TAKS is. I haven't been able to find the full report for Texas, but the summary (PDF) show that for many tests, a passing score on the TAKS puts the student around the 30th percentile nationally. For third grade reading, the apparent TAKS cut-off puts the child at the 12th percentile nationally.

Again, let me make it clear that I do think that there is a real role for (low) minimum standards which every child should meet. My proposal here isn't to raise those minimum standards, but to encourage parents, schools, politicians and children to focus on higher standards as well.

Monday, February 9, 2009

No child gets ahead

Before I begin my rant about the No Child Left Behind program, let me state clearly four views that I enthusiastically share with the program.
  1. Every child should become at least minimally proficient in certain essential skills
  2. It is import to provide a way of comparing how well schools and districts achieve educational goals
  3. Incentives to schools and districts do affect school policies.
  4. Standard tests are the least subjective way of measuring students' progress
Looking at that list, you might think that I was also an enthusiastic supporter of NCLB. But you would be wrong.
What the above leads to is a system where incentives to schools, school districts, property owners and so on is the portion of students in a school who meet a minimum standard. The overwhelming portion of what goes into a school's rating is whether a sufficient portion of the children reach a minimum standard.
Now let's look at the consequence of this
  • Effort going to serve the educational needs of children who will comfortably exceed the minimum standard will be wasted. That is, improvement among students who would meet the minimum standard in the first place will not be reflected in school ratings.
  • Effort going to serve students who are unlikely to meet the minimum standard is also wasted in that improvements there are not reflected in the ratings.
  • Success of a school, district or state becomes measured in terms of minimal standards, thus curriculum will be geared toward the minimum standard.
As mentioned in my points of agreement above, incentives do work. Therefore, we need to check very carefully what kind of incentives we establishes. All of the negative consequences that I've listed are the result of the incentive system that NCLB creates.
How not to fix the system.
The system that we have in place provides us (at best) with measures of how well schools have students meet minimum standards. The worst thing that we could do is to treat that measure as anything other than what it is. It is not a measure of how well the school educates students on average; it is not a measure in any real sense of how "good" a school is. The only thing that this measure tells us is whether or not a school is failing to get most of its students to meet a minimal standard.
But because this is the only measure that we have for comparing schools and districts within a state, it gets used for purposes that it was never designed to serve. A test of how many students meet a minimal standard can only be used to identify failing schools. It is entirely useless at distinguishing decent schools from excellent schools. For example, in a successful district like Plano Independent School District, more students hit the ceiling (reach the maximum score possible on the test) than actually fail it [get source for this. It was buried in one of the NWEA documents]. The Texas Assessment of Knowledge and Skills (TAKS) completely fails to provide a measure that can be used to compare those who well exceed its minimum standard.
A natural reaction and proposed solution is to raise the tests' standards: Make the test harder. But that would actually make matters worse. If we raise the standards of the test to a degree where it would be meaningful for schools that are doing well within the state, then it no longer works as ensuring that all students reach a minimum standard. It will place the higher standard out of reach of some students. The fact of the matter is that not all students are alike. Not all students are college bound. And when we establish truly minimum standards we need to take that into account. If our "minimum" is no longer truly minimum and so become out of reach for some children, then those children will be left behind.
Of course we should set high goals for everyone so that we get the best from each, but we shouldn't insist that everyone reach high goals.
Possible fixes
I have several ideas for how to improve the system. I will sketch a few of them below, but will need to expand on them at some later date.
Make every score count
Instead of reporting just the number of students who merely pass the test, also report the average score. This way an improvement for any student (whether well below passing or well above passing) gets credited to the school. A simple average may not be the best number because of how outliers affect the results, but some statistic or set of statistics that makes every child's performance count will help motivate schools to help each and every child whether they are near the passing threshold or not.
But this can't be done with the test as it stands (at least in Texas). As I mentioned, the Texas test has a very low ceiling. In my district more students score 100% on each test than fail it. This means that the test is providing absolutely no useful measurement for those students. Also when scores reach a ceiling they have a perverse affect on averages. Developing a test which has the appropriate scoring system is difficult, but achievable. But I will leave that for another post as well.
Distinguish between pedagogically useful and useless tests
Tests like the TAKS provide little information to the teacher about how to help an individual child. The tests are not pedagogically useful. That's fine because that is not their intent. They are designed to help us compare schools and districts. Unfortunately a great deal of school, teacher, and student effort goes into pedagogically useless activity. There are several ideas of how to deal with this, all involve less individual testing. One would be to test less frequently, and another would be to test only a sample of the students in a school instead of testing the entire school population.
Simplify administration of tests by eliminating the conflict of interest
The rules and procedures that schools have to follow for administering the tests are beyond belief. Visit your local school and ask to see the printed guidelines. You won't have time to read them, and you may not even have time to count the pages. Just weigh them on a bathroom scale.
The reason for many of these rules is because there is a truly awful design decision in the administration of the tests. The people who have the most at stake (the teachers and the school officials) are the ones who are asked to administer the test. This is a massive conflict of interest. And if you believe, as I do, that incentive systems do affect how people behave, then you see that there is a terrible opportunity for test administrators helping students cheat on the tests.
Of course I believe that most teachers are honest, but I also lock my car even though I think that most people wouldn't steal it. If you ask any teacher about cheating on these tests they will of course tell you that it doesn't happen in their school, but they will also be familiar with some terrible abuses in other schools or districts. We have set up a systematic administrative conflict of interest and so add boatloads of Band-Aids to cover a wound that is wider than a church door and deeper than a well.
The entire administration of these non-pedagogical exams could be simplified if we had them administered by some third party.
I will try to expand on these thoughts and fill in many of the blanks in further posts. But at this point, I should just upload this one even though it fall far short of what I had hoped for it.