ED 243 964 AUTHOR TITLE INSTITUTION PUB DATE NOTE AVAILABLE FROM PUB TYPE EDRS PRICE DESCRIPTORS IDENTIFIERS ABSTRACT DOCUMENT RESUME TM 840 287 Bliss, Leonard B. Item Analysis and Report of Student Skills of Secondary School Students. Technical Report #2, Public School Basic Skills Achievement Survey. - College of the Virgin Islands, St. Thomas. Caribbean Research Intt. Jun 82 103p.; For Technical Report No. 11 see ED 227 150. Caribbean Recaz.,.h Institute, College of the Virgin Islands, St. Thomas, Virgin Islandt 00801 ($4.00). Reports Research/Technical (143) MF01/PC05 Plus Postage. *Academic Achievement; *Achievement Tests; *Basic Skills; Difficulty Level; Elementary Secondary Education; *Item Analysis; Language Skills; Mathematics Skills; Reading Skins; *School Districts; Test Items Stanford Achievement Tests; Stanford Test of Academic Skills; *Virgin Islands A sample of slightly over 1,500 students from even-numbered grades in public schools of the U.S. Virgin Islands were given the 1973 edition of the Stanford Achievement Test (in grades 2, 4, 6, and 8) and the Test of Academic Skills (grades 10 and 12) in an attempt to assess student academic achievement in the basic skill areas of mathematics, reading, and English Language. This report describes the data analysis which involved a detailed item analysis of each item on each test given to the sample of students in grades 8, 10, and 12 as well as a summary of student skills based on their achievement along objectives provided by the test publisher and keyed to individual test items. Measures of item difficulty and item discrimination were calculated for the entire territorial school system and for the individual school districts. (Author) ***********************************************************************. * Reproductions supplied by EDRS are the best that can be made * * from the original document. * ********************************************************),************* Virgin Islands of the United State's Public School Basic Skills Achievement Survey Technical Report # 2 Item Analysis and Report of Student Skills of Secondary School Students Caribbean Research Ingtitute; College of the Virgin Islands Leonard B. Bliss, Ph.D. - Principal Investigator June 1982 Contents 1 Abstract 2 Introduction 3 Background 5 Item analysis 6 Difficulty indices 6 Discrithination indices 7 Summary of student skills 11 Availability of data 13 8th Gi.'ado (Advanced Level) 14 Vocabulary 16 Reading Comprehension 18 Mathematics Concepts 29 Mathematics Computations 32 mathortics Applications 36 Spelling 40 Language 43 10th Grade (TASK Level I) 51 Reading 53 English 63 Mathematics 67 12th Grade (TASK Level II) 73 Reading 75 English 84 Mathematics 88 Discussion 96 References 100 4 Abstract A sample of slightly over 1500 students from even numbered grades in public schools of the U.S. Virgin Islands were given the 1973 editiOn. of the Stanford Achievement Test (in grades 2, 4, 6, and 8) and the TeSt of Academic Skills (grades IQ and 12) in an attempt to assess student academic achievement in the.basic skill areas of mathematics, reading; and English Language. This report deSdribeS Phase II of the data analysis which involved a detailed item analysis of each item each test giv6n to the sample of students in grades ;--_ 8, 10; and 12 as well as a summary of student skills based on their achievement along objectives provided by the test publisher and keyed to individual test items; Measures. of item difficulty and item discrimination were calculated for the entire territorial Sdhbbl system and for the individual school districts At the request of the US. Virgin Islands Department of 8-du-cation and the iioard. of Trustees Of the College of the Virgin Islands, the Caribbean Research Institute embarked of basic skills achieVeMent i It soon became clear that virgin islands any Strategy designed to skills achievement needed to Start Off with on a study public schools, improve baSid _ \ a fairly detailCd description of current achieVeMent levels of student:s in territorial public schools It Was founa that this was not available. The task force set up to design the study decided that the Most efficient way to obtain information on levels of basic skills achievement was to administer a Standardized achievement test to a representative sample of students and to analyze the 1982) results of this test. Technical_ Repo_r_t_l (Bliss, describes, in detail, the process used to choose an appropriate test; Finally, the Stanford Achievement Test (1973 version) was chosen as the instrument of choice. Briefly, the reasons for choosino this adhieVetent test battery were that 1) it covered the grades K-12, 2) it seemed, on initial oCServation, to be a good match with subject content taught in the schools, 3) it was technically sound given the population on which it was standardiZed (a seemingly representative sample of mainland U.S. students); and 4) it would report out criterion referenced results; Due to various organizational and fiscal contraints only students in even numbered grades were tested. This seemed acceptable since many of the objectives tested by the Stanford Test carry across adjacent levels of the test an-, no reason to suspect that the patLerns of academic of students in odd numbered gra s were difforen in even numbered grades. These constraints are in Technical Report #1. of students to be tested was done using a technique with classes as the cluster unit. Ti-ii Ciet; 17=iplino procedure can be found in Technical lepOi-t discussion of the effects of using this sami, On the information obtained. Table I presents a awn of the sample as it finally emerged; Table 1 U.S. Virgin Islands Sample Sizes Test Level Total St.Thomas System St. John District St. C. 1)i st. r TASK II 129 74 TASK I 254 167 / Advanced 345 173 1 7 Intermediate II 227 146 Primary TIT 346 186 Primary I 234 143 ji Total 1535 889 (.16 3 'sting was done at the grade level recommended by the `c.. c. i-,UbliSher during the Week of October ni 1980 in the r:MiS/St. John district and dUting the week of December 1, in thO St. Croix district. Testing materials and completed -et.-0 collected; answer Sheets checked to deterMine with marking instructions; and answer sheets sent to ')o machine scored. BackgrbUnd - Technical Report 42 is the socond of four reports that Will deal with thc :_:osults or the basic skills assessment deSeribed above. Tc!chnil Report U. detailed the procedures used in test selection SaMpling; and test administration; More importantly; it estblished empitically; the content validity and the rilLibility of the Stanford AchieVeMent Test when-it was a,Thihistered to a Sample of b;S. Virgih Islands bdbli6 school This iS particularly important since there exists i:JcLr(lizod test Of academic achievement which includes t(idonts in its Standardization group. report examines the scores of secondary school students 0; and 12) and presents : I) al: item analysis of each item on each test of the battery which includes indices of item difficulty and discrimination and 2) a summary of student skills based on their scores on the itOM8 keyed to specific objectives. 6 ii-em _Analysis Dfficulty indite8 The difficulty index of an item is found by taking the number of examinees scoring correctly on an item and diViding it by the total number of students taking the test. In shOrt, it is the proportion of examinees who scored correctly on an item and has a range from 0 (no students SCoting correctly) to 1 (all students scoring correctly). BecauSe of the fact it is a proportion, it is often designated in the literature as "p" (e.g. p=.75). It may be worthwhile to point out that the term "difficulty indo" may be somewhat of a misnomer since items with high difficulty indices are actually leS8 difficult than items with low diffitulty indices-. NevertheleSS, the term and its definition have become standard in the area Of psycho- metrics throughout the United States. Difficulty indices for each item on each test were reported out by the test scoring service. In addition; difficulty indiC88 for examinees in the standardization group in the same grade as local examinees at approximately the same time of year are reported. The test scoring service used a Chi-squared test for proportions Oh each difficulty index to test the hypotheses that the proportions of local students scoring correctly on individual items in greater or less than the proportions of examinees in the standarditatiOn group scoring correctly at the .05 level of significande. Significant differences in either direction were reported out. 9 7 DiscriMinatien incicos The itern Oicrr1natLon :ndex indicates the degree to Which resp6nse:, Ch ltaM related to responSeS On Other items on he teeL. ThO itic indicates whether a person who does well on thc? tt Whole (that is a person who iS presumably high On the trait betricj measured) is more likely to -9iet the particular iern correct than a person who 6OeS poorly on the test as a whole. In other words; the item diSCriMinatic:i whethet ail item discriminates between thdSe who lo an those who do poorly on the test as a whele. Taking the item difficulty and the item discrimina- tion indek ibtb COnSideratioru the developerS of tests desire to construct tests WhiCh discriminate well among examinees with varying' levels of a trait, The item disc:rtMination index is calculated by the formula , L Where = thL:newiher Ohb have total test scores tOlal test scores and who also have the it.pr; cnrrct. L = the numb'el: of examineos Who have total test scores in the lbwy range of total test 'scores and who also have the item coi-ect N - the in the upper or low range C) 1. the test score,. 8 By definition; d is the difference between the proportion of high scoring examinees who got the item correct and the proportion Of loW scoring examinees who got the item correct; Upper and lower ranges generally are defined as the upper and lower 10% to 33% of the sample, With-examinees ordered on the basis of their total test score. When total scores are normally distributed, using the upper and lower 27% produces the best estimate of d (Kelly, 1939). If the distribution of total test scores is flatter than the normal curve, the optimum percentage is larger and approaches 33% (Cureton, 1957). However; Allen and Yen (1979) fOund that; for most applications, any percentage between 25 and 33 will yield similar estimates of d. In this study; 27% was used as the upper and lower percentage because examination of selected distributions Of actual test scores revealed nearly normal distributions. The theoretical range of d is between -1 and ±1. However; maximum discrimination is likely to occur when p=.50: When p=.50 the Varian-C-6 in item scores, which is p(1-p), is maximized. As an item becomes more difficult, it is less likely that any student will score correctly on it. As it becomeS less difficult it is more likely that any student will get it correct. This could lead to the suggestion that all items should have p=.50, but the useful- ness of this suggestion is mitigated by interrcorrelations among items. In an extreme case; if the items on a test all interrcorre- latod Perfectly and had difficulties of .50i half the examinees would receive a total test grade of zero and the other halt would have perfect test scores. Hence, there would be no fine discrimination between examinees' levels of achievement or whatever trait Was being measured; _., general, test designers tend to try to choose items with a range of diffitUltieS that average around .50. IteMs of particularly low diffietlty are often included in tests (usually among the earlier items) for motivational reasons. Itefii discrimination indices were calculated in this study to provide indications that items may be flawed when used with U VI students: Such flaws are ambiguity; the presence of clues; the presence of more than one correct answer, and other technical defects;- If none was found upon examination of the iteM, and it was determined that the item did; indeed; appear to measure the objective it was intended to; the i-6-2:M was indlUded in the overall analysis of results. Any item that discriminate positively can make a contribution to the measure- ment of pupil achievement and low indices Of discrimination are frequently Obtained for reasons other than iteM defects; Standardized achievement tests are designed to measure several different types of learning outcomes (e.g. knowledge; understanding; application, etc.). Where this is the case, test items that represent an area receiving relatively little emphasis will tend to have pbbr discriminating POw6r. For example, if a test has forty item measuring kh6isi16d4e of specific facts and ten items measuring understanding; the latter items can be expected to have low disatimination indices. 12 This is because the items measuring understandihg have less representation in the total test score and there is typically .a low correlation between measures of knowledge and measures of unde-standihg. LOW discrimination indices here merely irldicate that these items are measuring something different_ from what the major part of the test is measuring; Removing such items from the test would make it a more homogenous measure of knowledge outcomes; but it would also damage the content validity of the test because it would no longer measure objectives in the understandihg area. Since achievement test batteries need to measure a wide variety of objectives in a reasonably short period of time; they tend to be fairly heterogeneous in nature and moderately low discrimination indices tend to be the rule rather than the exception. To summarize, a low discrimination index alerts test users to the possible presence of defects in test items but doeS not cause them to discard these items if they appear to be funetiOh- ing as they should. A well constructed achievement test will - of necessity; contain items with low discriminating power and to discrard them would result in a test which is less; rather than more; valid; Dthe to these considerations, in this study items were examined if they had diSeriminz:tion indices lower thah .20. This is a rather conservative criterion since items that diSCtiminate as low as this may provide useful information; but given the unknown test taking charactcristics of USVI students; it was decided to be particularly cautious in the item analysis. 13 11 Summary of Student Skills The items on the tests of the various levels of the Stanford AdhieVetent Test battery were keyeii3 tti behaviorally stated instructional Objectives. These objectives are grouped into a series Of item groups.. Tables are aVailable which present objectives within item groups and the difficulty and discrimination ildiceS for the item or items which evaluate those objectives; TheSe may be obtained from the Caribbean Research Institute; In addition, the difficulty index for each item for the examinees in the Standardization group is available. This standardization group consists of examinees who were in the same grade at approximately the same time during the school year as the U.S. Virgin Island8 sample; The national p values are used not as a means to compare U-S. Vitain Islands students with mainland US examinees. Historital and cultural differences ti-otweon these two groups of examinees makes this comparison an inappropriate one. Philosophital tonsiderations aside, hoticievet, such comparisons are of little use to the people who make curricular decisions in schools. What these people need to knOW are the particular levels of skills of students as measured aqaitist well defined objectives, not hOW well their students achieved thse skills as compared to other students. Nevertheless, since the skills and knowledges taught in schools are seldoM taught once, but are dealt with at a number of grade levels where they are reinforced and broadened, the level of achievement= on specific obje tives should_be expected 12 to change from grade to grade for a particular student or group of students. The publishers of the Stanford Achievement Test take this factor into consideration by testing particular objectives across a number of levels (grades) of the test; this study the national p valuea are used to indicate the relative level of difficulty of an item by which the performance of the local sample nay be judged For instance, it would be foolish to be dissatistified if 20% of the USVI students have indicated that they have reached an objective when only 18% of comparable students in the standardization group had reached that same objective. What is more likely the case is that this is a difficult and complex objective .t.hat had just been taught recently and would be retaught and elarged upon at a later time. Thetefbte; the following criteria were used in summarizing student skills. Skills are described as "adequate" if the propor- .tions of local examinees scoring correctly on items measuring those skills are not significantly higher or lower than the proportions of the standardization group scoring correctly or, if significantly higher or lower; the proportions correct are within 10% of the standardization group proportion correct as reported by the test scoring service; Skills are described as "strong" if the proportions of local examinees scoring correctly on items measuring those skills are significantly higher and more than 10% greater than the standardization group proportion correct as reported by the test scoring are described, a8 "weak", if the proportion of 15-dal examinees scoring cOrrectly'on items measuring those skills are 13 significantly 16Wet and more than 10% less than the standiza- tion group proportion correct. The 10% standard was set*in the realization that some differences, while statistically significant may be educationally trivial and it was noted that most differences indicated as significant exceeded 10 %. Findlly, in cases where the scores of examinees from both the St. Thomas/St; John and the St. Cr6iX districts were the same based OA the criteria stated above summaries were based on the entire USVI sample; Where differences were noted, the skills of e:-