ED 227 150 AUTHOR TITLE INSTITUTION PUB DATE NOTE AVAILABLE FROM PUB TYPE EDRS PRICE DESCRIPTORS IDENTIFIERS ABSTRACT DOCUMENT RESUME T 830 161 Bliss, Leonard B. Validation of the Use of the Stanford Achievement Test with U.S.V.I. Students. Virgin Islands of the United States Public School Basic Skills Assessment Survey, Technical Report No. l. College of the Virgin Islands, St. Research Inst. Jan 82 54p.; Paper presented at the Annual Meeting of the American Educational Research Association (66th, New York, NY, March 19-23, 1982). Caribbean Research Institute, College of the Virgin Islands, St. Thomas, USVI 00802 ($2.00). Speeches/Conference Papers (150) -- Reports - Research/Technical (143) Thomas. Caribbean MFO1/PCO3 Plus Postage. *Achievement Tests; *Basic Skills; Data Analysis; Educational Assessmeni; Elementary Secondary Education; Language Skills; Mathematics Achievement; Reading Achievement: Sampling; *Standardized Tests; *State Programs; *Test Reliability; *Test Validity *Stanford Achievement Tests; Virgin Islands A sample of slightly over 1500 students was drawn from even-numbered grades in public schools of the U.S. Virgin Islands, and was given the 1973 edition of the Stanford Achievement Test (in grades 2,4,6, & 8) and the Test of Academic Skills (grades 10 and 12) to assess student academic achievement in the basic skill areas of mathematics, reading, and English language. This report describes phase I of the data analysis, which involved the determination of levels of content validity and reliability of the scores obtained from these Virgin Islands students on these tests which were originally standardized on continental United States populations. The results indicate that the tests are content valid for use in Virgin Islands public schools at these grade levels and that the scores obtained are at least as reliable as those obtained using continental U.S. students during the test standardization procedures. (Author/PN) REET KEKEKEKEREKEKEKEEEEKEEREKREKEREKEKEEKEEKKEREKKEEEKEEKEEREREEKEKEKEEEKEKER ® Reproductions supplied by EDRS are the best that can be made " * from the original document. e® KRAEEEKEEKEKEEKEKEKEEKEEKEEKEEEEKEKEKERERKEREREKEKKEKEREEEREEEEEEKEKEEEERREEEREREER U.S. DEPARTMENT OF EDUCATION NATIONAL INSTITUTE OF EDUCATION EDUCATIONAL RESOURCES INFORMATION CENTER (ERIC XK This document has beer reproduced as the person or organization ED227150 Virgin Islands of the United States Public School Basic Skills Achievement Survey Technical Report #1: Validation of the Use of the Stanford Achievement Test With U.S.V.I. Students “PERMISSION TO REPRODUCE THIS MATERIAL HAS BEEN GRANTED BY L.0 3) rss TO THE EDUCATIONAL RESOURCES INFORMATION CENTER (ERIC).” PRESENTED AT THE ANNUAL MEETING QF THE AWERICAN EDUCATIONAL RESEARCH ASSOCTATION NEW YORK CITY MARCH I9-235, 1982 / Caribbean Research Institute, College of the Virgin Islands Leonard B. Bliss, Ph.D. - Principal Investigator January 1982 PREFACE With the appearance of Virgin Islands of the United States Public School Basic Skills Achievement Survey, Tech- nical Report #1: Validation of the Use of the Stanford Achievement Test With U.S.V.I. Students the Institute has embarked on the publication of a Working Paner Series. These papers are intended to present the author's (and the Institute's) point of view on various subjects as a matter for discussion and comment by those who agree as well as disagree with expressed positions. In this way the Insti- tute hopes that the final versions will be improvea in style as well as rigour. The present paper is the first phase of a study of basic skills in the schools of the United States Virgin Islands requested by the Board of Trustees of the College. The work has taken considerably longer than anticipated due to fundamental alterations in the design so as to provide greater depth than originally planned. Unfortunately, shortage of staff did not allow the progress hoped for to be made. The data for the whole project have been collected, however, and work is proceeding on their interpretation and the compiling of the three reports which will follow. Norwell Harrigan Director Abstract A sample of slightly over 1500 was drawn from even numbered grades in public schools of the U.S. Virgin Islands and were given the 1973 edition of the Stanford Achievement Test (in grades 2,4,6, & 8) and the Test of Academic Skills (grades 10 and 12) in an attempt to assess student academic achievement in the basic skill areas of mathematics, reading, and English language. This report describes Phase I of the data analysis which involved the determination of levels of content validity and reliability of the scores obtained from these Virgin Islands students on these tests which were originally standardized on continental United States populations. The results indicate that the tests are content valid for use in Virgin Islands public schools at all of these grade levels and that the scores obtained ave at least as reliable as those obtained using continental U.5. students during the test standardization procedures. It is almost becoming a matter of faith that achievement in basic skills (i.e. English language and mathematics) in public schools under the American flag has deteriorated over the last twenty years. Proponents of this idea point to evi- dence as formal as decreases in typical scores on the Scholas- tic Aptitude Test and standardized tests of academic achieve- ment and as informal as the quality of writing and arithmetic skills they perceive in the young people around them. The reactions of people to this perceived phenomenon are also varied. On the goverament level they include the require- ment that all students score a minimum grade on a test of basic skills in order to receive a high school diploma; that teachers pass a similar test to obtain teacher certification; and that schools require students to take additional course work in basic skills areas. In addition, federal, state, and local governments have initiated programs to provide support in the forms of grants and technical assistance to schools at all levels to do research and set up programs designed to improve student achievement in basic skills. At a different level, parents, concerned that the public schools are not doing an adequate job in preparing their children in basic skills areas, are choosing, in increasing numbers, to remove their children from public schools and place them in religious and secular private schools. While there are other ~ 2- reasons for the proliferation of private schools besides the purely acadeinic, the desire for high quality academic prepara- tion is one compelling cause of this phenomenon. The public schools, themselves, have reacted strongly to this crisis in public confidence. These reactions include an increase in required courses in language and mathematics areas with a corresponding decrease in electives in areas considered less "basic." Projects to revise curricula in basic skills areas proliferate and are receiving more support than they have Since the reevaluation of American education engendered by the shock of Sputnik in the late 1950's. Improving basic skills achievement was a concern of the Department of Education of the government of the Virgin Islands of the United States when it approached the College of the Virgin Islands to provide aid in improving such instruction. In an effort to provide this service, the Caribbean Research Insti- tue, the college's research arm, worked with a task force com- posed of representatives from the Department of Education and CRI to determine a course of action. It became clear after the first few task force meetings that development of any strategy designed to improve basic Skills achievement needed to start off with a fairly detailed description of current ac ievement levels of students in terri- torial public schools. This information was not available. Public school students were administered a standardized achievement test only at the end of sixth grade (The Iowa Test of Basic Skills). In other elementary grades most students ig, YEP were tested annually or semiannually at their schools, but the test given and the times during the academic year that were administered varied greatly and apparently at the whim of build- ing administrators. The results of these tests stayed at the schools and were not collected at any central point. On the secondary level there was no program of standardized achieve- ment testing. An additional factor which limited the use of previously collected achievement level cata was that all scores were re- ported in a norm referenced manner. That is, scores did not adicate which basic skills examinees had or lacked, but rather now examinee's scores compared to those obtained by a group of students to whom the tests were previousiy administered in the continental United States. The Iowa Test of Basic Skills administered to sixth graders did make comparisons with other V.I. sixth grade students (i.e. they reported using local norms), but even these were of no use in determining whether or not individual students had attained specific basic skills. It was decided to test a representative sample of U.S. Virgin Islands public school students using a standardized basic skills battery. Choosing the test, the following criteria were used: 1) The test must be technically sound in terms of reliability and item discrimination, at least for the group it had been field tested on. The test must be content valid for U.S. Virgin Islands public school students. That is, there needed to be a high degree of matching between the content and behaviors sampled by the test and those actually in the curriculum taught at various levels in the U.S.V.I. public schools. The test must include a detailed statement of the objectives tested while providing an item by objective keying procedure. Scores which indicate students' performances relative to each objective must be available. That is, criterion-referenced scoring must be provided. The 1973 version of the Stanford Achievement Test (Basic Battery) was chosen as the test which appeared to meet the criteria listed above. It was administered to slightly over 1500 students in the Fall of 1980 in both the St. Thomas/St. John and the St. Croix school districts. This is the first of a series of research reports designed to make available the results of this rather complex study. A simple, brief example of the quancity of data obtained may serve to highlight the scope of this study. The Intermediate Level II of the Stanford Achievement Test (administered to sixth graders in this study) contained 351 items. It was administered to 225 students in the U.S. Virgin Islands sample yielding 78,975 individual pieces of data. The sixth grade sample, due to a technical difficulty (the principal in one school forgot to assign the teacher of the selected class the task of giving the test and the teacher in another school administered only four of the sever. subtests), contained the smallest number of examinees of any grade level. Additional reports will be issued regu- larly as soon — X oO x > +e >m.o be cderined — a A ress rm" ry Ww ARES ae op Ae ws | oOo onwaso dy pete a>) / . fy yr? fy Oo r So woe Ae) jon fo =| aa) measuring ti I J it Since this study is operationally defining basic skills achievement as the performance of students on the Stanford Achievement Test, it is clear that no construct is being measured. Hence, construct validity will not be a concern in this report. The content of any curriculum can be thought of as being composed of »uvject matter content and behavioral! changes sought in students. For a test to be content valid it must provide re- sults that are representative of the topics and behaviors we wish to measure. More formally, ". . . content validity may be defined as the exter: to which a test measures a representative sample of the subject matter and the behaviorez1 changes under consideration” (Gronlund, 1976, pp. 81-82). Effective strategies for determining content validity involve determining the objec- ! — f£ , ) e rer ~fr, _~ ‘ =. - degree of match betwee . AcnhLevement are primarily concerned with measuri cquisition of cer- tain skills and knowledges (objectives) by students at the time that the St is given. it Ls content validity that ’ ’ 7 . . ~* snould be of prime conc pecil the ob lectives —_ 4 his give highly reliable scores for one group examinees, but result in lower reliability with another group. [In essence, what we are concerned with is whether or not the test scores represent measures of the same traits each time the test is given. important, then, that whatever measure of basic used, that the measure be content valid schools and 7 Islands -9- (1975) indicates that a sample of over 275,000 pupils from 109 school systems in 43 states in the United States made up the standardization samples used. Table 1 provides descriptions of these samples and how they compare with a description of the population of the continental United States. Content validity was established by curricular analysis using information from a large number of sources. Basic to the construction of a series of achievement tests is the identification of what is being taught in the scnools across the nation. The most important sources for curricular analysis were (a) textbook series in various subject areas (including the prepa- ration of detailed analysis of the content of the books most widely used in each field); (b) a wide variety of courses of study from individual school Systems; (c) statements of objectives from various state and national committees, and the opinions of experts in various fields; and (d) the research literature pertaining to children's concepts, expe- rience, and vocabulary. (Technical Data Report, p.12) The reliability of the scores of the standardization sample was determined by using the Kuder-Richardson Formula 20 and by calculating the standard error of measurement of the scores. Two measures of reliability were used since it is known that high homogeneity in tested groups will lower the reliability estimates obtained using the KR-20, but that his effect is dealt with in determining standard errors. In addition, the standard error of measurement is more meaningful in interpreting scores of individual students. With very few exceptions, the reliabil- ities obtained from the standardization samples ranged from .84 to .95 using the KR-20 formula. While the 1973 version of the Standford Achievement Test appears to be educationally sound based on the standardization Table 1 Summary of Characteristics of Standardization Samples | National Stanford Stanford U.S. Population Characteristics Population Range 1970 Data i iat earner tein een, pe A tern areas i eteietendntmsmenipnsitiniaasce Percent of pupils by community size 0-49, 999 50,000-249, 999 250,000 or more Percent of pupils by Geographic Region Southeast North Central Northeast West Median Family Income $ 4,878 to $13,593 Median Years of Schooling (Adults 25 yrs. & older) Average Class Size (Student-Teacher Ratio) 36 Average Starting Salary of Teachers $7,116 $ 4,500 to $7,064 11,500 Average Salary of Teachers $9,360 $ 4,500 to $9,265 11,500 Median Years Teaching Experience ‘ 5 to 24 Percent of Grade 1] pupils who attended kindergarten ‘ 0 to 100 Percent of Schools Using Some Team Teaching Table 1 continued Stantord Population Characteristics Percent of Schools Using Some Teacher Aids Percent of Pupils Not Promoted to Next Highest Grade Grade Grade Grade Grade Grade Grade Grade Grade Grade L > 3 4 5 6 7 8 9 Percent of Pupils ’ Non-public Schools Percent Ethnic of Major Minorities Blacks Hispanics Other Stanford Range = =e OO & & Whe OwoWm Ww rvwo ho Pe 3240 4.6 Less than lL National U.S. 1970 Data sa a | 4.6 Less than l Population lrrom Stanford Achievement Test: Technical Data Report, p. ys groups data, the groups contained only continental U.S. Students. Likewise, the test makers most probably did not take Virgin Islands public school curriculum into account when designing items. Therefore, before the scores of any tests of basic skills can be used to draw conclusions about V.1. students, the content validity and reliability of these test scores for Virgin Islands students must be established. Hence, this report. Samp Ling The June 1, 1979 enrollment in the public schools in the Virgin Islands of the United States was 25,426 according to the statistics issued by the V.I. Department of Education. It was clear that testing this number of students was economically unfeasible. The preferred alternative would have been to generate a random sample of students in grades K-12 to be tested, but it was equally clear that this would have produced an in- tolerable disruption of classroom activities. Therefore, in an attempt to obtain a representative sample of s adents, cluster sampling was used with the clusters being defined as classes. The number of classes to be selected for the sample from each grade in eacn of the St. Thomas/St. Jonn and St. Croix districts was determined by calculating the proportion of the total K-12 student popuiation in each grade in each district and assuming a class size of thirty. Selecting whole classes presented an additional difficulty. The smail number of classes selected in each grade might have made obtaining a representative sample of students more diffi- cult. This is due to the fact that while classes in a given elementary school may be heterogeneous, the schools themselves are not. This is because elementary schools in the U.S. Virgin Islands are essentiaily neighborhood schools. Virgina Islands neighborhoods tend to be homogeneous in terms of socioeconomic status of residents. To overcome this problem, it was decided to increase the number of classes tested in a given grade «Then (thereby increasing the number of schools within the territory from which these classes came) without increasing the total number of students tested by testing at alternave grades. This seemed acceptable since many of the objectives tested by the Stanford Achievement Test Carry across adjacent levels of the test and there was no reason CO suspect that the patterns of academic achievement of students in odd numbered grades were different from those in even numbered grades. [t was originally Proposed that students in odd numbered grades be tested during the Spring of 1980, but difficulties in obtaining testing materials resulted in testing being post- poned until the Fall of 1980. In order to deal with the cohort of students originally selected, even numbered grades were actually tested. The classes to be tested were chosen by chance. Specifi- cally, Zor each grade in each district a listing of classes was made and each class was assigned a number. A table of random numbers was consulted. Numbers were drawn from the table until there were the same number of random numbers chosen as there were classes needed for the sample. In the case of duplicate numbers being drawn, the duplicate was ignored and another number chosen. If the number chosen was outside the range of the number of classes on the list, it was ignored and another number was chosen. When sufficient numbers had been drawn, the listed classes which corresponded to these numbers were included in the sample. This procedure was repeated for each grade in each district. ud 35, The sole exception to this procedure was in the eighth grade portion of the sample. On St. Thomas, homeroom classes are somewhat homogeneous in that students repeating eighth grade and those in the eighth grade for the first time are placed in separate homeroom classes. Since the levels of academic achievement for repeaters and nonrepeaters are very likely different, the proportion of repeaters and nonrepeaters was determined to come out with a number o° classes needed in sample from each group and the groups of classes of repeat- and nonrepeaters were sampled separately in the mainer described in the preceding paragraph. Elena Christian Junior High School on St. Croix is on split session. The principal or that school fe'’t that there were definite differences in achievement levels between the students in the morning and afternoon sessions. Because of this, classes in the morning and afternoon sessions were sampled separately using the same procedure employed on St. Thomas for the repeating and nonrepeating homeroom classes. If simple random sampling has been used in selecting students to be tested, a sample size of approximately 2000 would have been the maximum size required to obtain an accuracy of about +2Z at a .95 level of confidence when estimating the proportion of V.I. students reaching certain objectives from the sample proportions if a typical proportion answering each item correctly were .50 (see Asher, 1979, p.166). In actuality, due to student absences, failure of school personnel to carry out requested tasks, and other difficulties, the sample size aus obtained was only 1535. However, examination of the difficulty indexes of the Stanford Achievement Test items on all levels revealed difficulty indexes considerably different from .50 On most items. This would tend to shorten the size of the confidence interval. Finally, the financial and organizational constraints cited previously forced the investigators to use cluster sampling techniques rather than random sampling. Since the intraclass correlations (i.e. the effects of clustering on the standard deviations of the achievement test scores) were not known, this factor also contributes toward making the above mentioned accuracy estimate a rather crude one. It can, however, serve as a rough guideline. Table 2 preserts the relevant sample size data. The sixth and second grade samples from St. Croix are smaller than had been hoped for the following reasons. As indicated pre- viously, the ceacher of one of the sixth grade classes only administered four of the seven subtests. In a second grade class, the teacher was ill during the days set aside for test- ing and the test was not administered. By the time this became apparent to the investigators, it was too late to go back to St. Croix to retest Aside from the difficulty in estimating precision of the Proportions of students obtaining correct scores on various items, the sampling procedure used presents another difficulty. Because of the previously stated practical considerations, it was necessary to employ cluster sampling (sampling whole classes) rather than simple random Sampling o- students to be tested. rable Virgin Islands Sample Sizes Total St.Tnomas/St.John St.Croix Grade eV System District District TASK II 74 55 TASK | Z 167 Advanced 345 173 Intermediate II a 146 Primary III 186 Primary ne 143 889 [he principal drawback to ciuster sampling is the likelihood of increased sampling error. In general, as the size of the sample increases, the size of the standard error decreases. This applies, however when each sample element (in this case, each student) is Se Legted independently o1 every other element. In cluster san- pling the elements are, by definition, selected in a group rather than indépendently. The effect of clustered selection on the standard error will depend on the similarity between the elements inthe c.uster and tnose in the population. In many cases, sample elements selected in clusters will not show the same variation as an equivalent number seiected independently. Students who attend the same school and are in the same class may be more like one another in a characteristic such as aca- demic achievement than students in the public school population as a whole. The relationship between clustering and sampling error may be summarized as follows. If all the elements (students) in a w TBs Cluster (class) were identical with regard to achievement and totally different from the elements in other clusters, the sampling error would be extremely high. Clustering, inthis case, would tend to make the clustered sample equivalent in Size to a simple random sample with as many subjects as there are clusters, rather than elements. Hence, a sample made up of 60 clusters might be equivalent to a simple random sample of 60 individuals. This is obviously an extreme case that is never seen in practice. At the opposite extreme would be a series of clusters showing the same variation within each cluster as simple random samples of the same size. In this case, each cluster would represent the entire population, another condition rarely met in practice. Most sampling situa- tions fall in between these extremes, tending toward one or the other according to the characteristic being studied. In general, according to Warwick and Lininger (1975), experience has shown that well-designed cluster samples will produce standard errors that are about one and one-half times as large as the standard errors from simple random samples of the ho lon) 2) hs 5 a es 6. Zs 2. be mw Grade 6-Intermediate II Level 22. 31. 29. 19. 24. 19. ‘> af — eh ee SFnNLfON AUD Ww ow We bo DBaAolfnre LO OUMNH Owl won — Grade 4-Primary III Level ya 42. 30. i. 13. 30. 28. Nm We hm & fh & ho COW CO WO dh Ww Cowl UW wow a me sell oon oounw } COM OF WwW OW ho KF f&onoaoanh £ Mean Stand. OnNnCOWM WSs! & Ceouwunlf OO, WY mMNmoUe & hb Ww wo Mean Stand ON OL RRP USA cCouNnrwronwno STX WIM C—O WU Ww Wo SFnNOON@WD@®DO Ss ER a System STT/STJ STX Mean Stand. Dev. Mean Stand. Dev. Mean Stand. ets tesensiatenemies een eeinesnesietinsiecerentes-- ee sare eeeinneettinheedinseatinaastinienisiannaunins a Grade 2-Primary | Level a Ae eeeneteeeemesen Vocabulary ; ~ Ke Reading (Part A) 34. ’ 36. Reading (Part B) ‘ : 30. Word Study Skills : . 48. Mathematics Concepts ‘ , 19. Mathematics Computation 21. ‘ 22. Listening Comprehension ‘ ’ Ads OSODwWnN ENO Teachers who administered the tests in their class- review the test publisher's ‘termine the degree of match ives and the basic skills they expected their sti a4 have obtained. Using these techni he researchers were satisfied the test id, indeed, test a sample ol objectives that consistent with the objectives used in teaching in the public schools of the Virgin Islands of the United States. Reliability The estimates of reliability of the test scores are pre- sented in Table 5. The KR-20 reliability estimate’ for each test is reported along with the KR-20 estimate for the mainland standardization samples as presented in the Technical Data Report. The issue of interpreting these reliability estimates is a complex one and will be dealt with in more detail at the conclusion of this report. The author felt the need to have at least a tentative criterion for making decisions regarding the accept- ability of the reliability estimates obtained from the V.I. sample of examinees. The Stanford Achievement Test is considered Yyx = [(n/(n-1)] [o y-Upq/(o y)]) where r,,= the reliability estimate (From Stanford Achieve- n= the number of scores ment Test: Technical o*y = the variance of the distribution Data Report, p. 35) of scores : p= the proportion of examinees marking the correct answer on a particular item q= isp -.6- to have more than acceptable reliability when administered to the population of examinees upon which it was standardized (i.e. continental U.S. students). Among the indications of this are numerous reviews of the test in the literature (Kasdon, 1974; Lehmann, 1975; Chase, 1978; Ebel, 1978; Thorndike, 1978) and the fact that it is widely used in the schools. However, the literature is replete with studies which indicate that standardized tests of academic achievement tend to produce less reliable scores when administered to students from low socioeconomic status homes and to those who are culturally different from the majority of those on whom the test was normed (see reviews and discussions in Anastasi, 1958; Tyler, 1955; and Deutsch, 1960). Therefore, if the reliability estimates obtained from a sample of U.S. Virgin Islands students who took the Stanford Achievement Test are at least equal to the reliability estimates obtained from the standardization samples, it is reasonable to conclude that the test scores are reliable indicators of academic achievemeat for these students. For each reliability estimate obtained from the V.I. sample, a reliability difference was found by subtracting the standard- ization groups' reliability estimate from the local groups' reliability estimates. The distribution of these differences is shown by the histogram in Figure 1. The median reliability difference was -.038 with a range from -.20 to +.05 with the distribution skewed to the left (i.e. negatively) quite markedly. In addition in an attempt to observe these reliability differences from another perspective, for each pair of relia- bility estimates (the standardization group estimate and the Table 5 Stanford Achievement Test Raw Score Reliability Estimates STAND. USVI ST THOMAS / GROUPS SYSTEM ST JOHN ST CROIX KR-20 KR-20 KR-20 KR-20 Grade 12-TASK II Level Reading Mathematics English Reading Mathematics English Grade 8-Advanced Level Vocabulary 89 | Reading Conprehension Mathematics Concepts Mathematics Computation Mathematics Application Spelling Language Grade 6- Intermediate II Level Vocabulary Reading Comprehension Word Study Skills Mathematics Concepts Mathematics Computation Mathematics Application Spelling Language Grade 4-Primary III Level Vocabulary Reading Word Study Skills Mathematics Concepts Mathematics Computation Mathematics Application Spelling Language Table 5 (cont. ) STAND. GROUPS KR-20 USVI ST THOMAS / SYSTEM ST JOKN KR-20 KR-20 ST CROIX KR-20 Grade 2-Primary I Level Vocabulary . 86 Reading Part A ~94 Reading Part B .95 Word Study Skills .93 Mathematics Concepts .81 Mathe..atics Computation .87 Listening Comprehension .77 .72* *Significantly lower than the standardization groups KR-20 at p=.05 ‘stimate), the hypothesis differences in tained were less than tested. Reliability imates wereé ranstormed using formations to normalize skewness the distribution correlation measures sted with One tailed sienifi indicated liability ion is in order in interpreting the results slgniticant differences As previously sampling was used in obtaining the sample students to be tested rather than simple random The res this is that the actual standard error le may very well be larger than the one used in cal- Statistic (estimated by [1/(N-3)) 21). The result of this would be that the values of ined were larger than they should have been and that some of the differences from -o that were noted in Table 5 to be Significant at the p=.05 level may actually not have been. put it in technical terms, the probability of Type I error is probably greater than .05 in each of This is a definite these hypothesis tests. in any conclusions we might draw from from a practical point of view, given the (Hayes, these weakness tests. However, decisions te be made, 1973, pp. 662-667) Figure l Frequency Distribution of Differences Between the Standardization Group Reliability Estimates and the V.I. Sample Reliability Estimates (Ar) the error of prefer stakenly assuming that scores are be that we would either look more closely at scores can yr discard the the testing as being unreliable for V.I. student S. lost is muc ime at poss y, some money. errors (mistakenly assuming th: sal scores are ¢ as reliable as the standardization groups' scores) would » to go anead and use the unreliable scores to make decisions about basic skills students and, possibly, to make decisions regard- l strategies that will be used in the schools. , what res 3 1s a rather liberal test of the hypotheses and, given the nature of the decisions to be made, this may not be totally undesirable. However, it must be kept in mind when interpreting these results that the actual level of Type I error is not known and that it is probably higher than +35. In any event, we can use Table 5 to flag tests where reliabilities may be less than acceptable. It was noted that, in the majority of cases, the variinces ox the raw scores obtained by the V.I1. sample were considerably lower than those reported for the standardization groups. This homogeneity is a phenomenon commonly found when testing samples drawn from populations composed largely of persons from low socioeconomic status homes. "The reliability of any test is partially dependent on the sample of individuals tested to obtain the coefficient. In general, the more heterogeneous the sample with respect to whatever the test is measuring, the higher the reliability coefficient will be" (Techaical Data Report), p. 35). The standardization groups' reliability co- efficients can be adjusted for homogeneity using the variances obtained from the local sample making reliability comparisons more meaningful. 4 Using the adjusted reliability estimates for the standard- ization groups, the differences between this group's and the V.1. sample's reliability estimates was calculated employing the same procedure used with the unadjusted estimates. The distribution of these differences is shown in the histogram in Figure 2. The median reliability difference using the adjusted estimates was -.002 with a range from -.06 to .02. As with the unadjusted scores, the distribution is negatively skewed, but not as markedly as with the unadjusted reliability estimate differences. When the standardization group's reliability estimates are adjusted for homogeneity, the differences between the reliability estimates of the two samples become fewer and smaller. Table 6 presents the adjusted estimates of reliability for the standardization groups' and the results of tests of the hypotheses that the differences between the standardization groups’ reliability estimates and the estimates of reliability | he for ] = ee 2 ws Ky Xk Using the formula Pry (1-6 i. /o =? (1 i ie where Pyy ando“, are, in this case, the reliability coefficient and variance of the standardization groups and px*y* and o2.* are the same statistics for the V.1. sample. (Hayes, 1973) ) Frequency Distribution of Differences Between th? standardization Group Adjusted Reliability and the V.I1. Sample A Reliability Estimates (Ar) Table 6 Adjusted Stanford Achievement Test Raw Score Reliability Estimates __USVI SYSTEM ST THOMAS/ST JOHN ST. CROIX Adj. Stand. Local Adj. Stand. Local Adj. Stand. Local Groups Sample Groups Sample Groups Sample KR- 20 KR-20 KR-20 KR-20 KR-20 KR-20 Grade 12 - TASK II Level Reading .93 91 Mathematics ; .87 .85 -nglish ; 91 91 Grade 10 - TASK I English Table 6 (cont. ) USVI SYSTEM ST THOMAS/ST JOHN ST. CROIX Stand. Local Adj. Stand. Local Adj. Stand. Local Sample Groups Sample Groups Sample KR-20 KR-20 KR-20 KR-20 KR-20 Grade 8 - Advanced Level .8] ae -35 ~ 9D . 80 .33 ics Conce pts matics Computation thematics Application spelling Language 44 Table 6 (cont.) USVI SYSTEM ST THOMAS/ST JOHN ST. CROIX Adj. Stand. Local Adj. Stand. Local Adj. Stand. Local Groups Samp le Groups Sample Groups Sample * KR-20 KR-20 KR-20 KR-20 KR-20 KR-20 Grade 4 - Primary III Vocavulary ' . 83 76 Reading Comprehension 92 eS Word Study Skills : .90 .89 Mathematics Concepts ; BY i .68 Matnematics Computation . 83 ‘lathematics Application .f . 86 .385 Spelling oa 91 Language . 86 . 84 Grade 2 - Primary I Level Vocabulary .76 ate a Sg | Reading - Part A 2 .98 i: .99 Reading - Part B eS mF «92 .91 Work Study Skills .91 92 .89 .91 Mathematics Concepts Be a .70 69 Mathematics Computation . 80 .81 .76 te Listening Computation _ .78 74 72 .70 *Significantly lower than the standardization groups’ KR-20 at p=.05 1 , obtained by n mentioned in di: s the hypothesis holds true. The in these tests comparisons showed Sinee the ifferences could be expecte as a result of chance). ferences significantly less d on a chance basis. error of measurement? for the raw scores of shown in Table 7 .when the reliability rpreted in terms of the standard error of problem of the influence of heterogeneity is taken into account, since the formula for of measurement includes the standard (Technical Data Baport, p. 35}. The measurement can be thought fF as the stan- lard deviation of ne ditferences between the scores obtained on the test and the true scores (the scores the examinees would have received if 1 ‘St were perfectly reliable). As such, it can be ‘d to determine an interval within which we.can be confident that the true score falls. For instance, we can be confiden that the rue score would be within one standard s the standard error of measure- the standard deviation of the r is the reliability coefficient. (Gronlund, 1976) Table 7 Stanford Achievement Test Standard Error of Measurement Estimates STAND. USVI ST THOMAS / os GROUPS SYSTEM ST JOHN ST CROIX Tesi S.E.M. S.E.M. S.E.M. S.E.M. Grade 12 TASK II Level Reading 2.60 3.65% Mathematics .80 : € English oe 3. Grade 10 - TASK I Reading Mathematics English 1 — + € Grade 8 - Advanced Leve Vocabulary 3.10 as Reading Comprehension 3.60 J. Mathematics Concepts 2.60 Mathematics Computation 2.90 Mathematics Application 2.60 Spelling 3.30 Language .90 THOMAS / JOHN ST CROIX Ewe, Intermediate Computatil \opoplic Grade 4 - Primary III Level Vocabulary Ye Reading Com Spelling Language Table 7 (cont.) STAND. GROUPS S.E.M. USVI SYSTEM S.E.M. ST THOMAS / ST JOHN ST CROIX S.E.M. S.E.M. Primary I Level Vocabulary Reading - Part A Reading - Part B word Study Skills Mathematics Concepts Mathematics Computation Listening Comprehension . 50 2.50 -40 . 80 2.30 2.20 2.0 2.68 re 1.95 : ee x ref .68 .65 , .68 . 38 ; 1 ae -i7 ei ay & | <22* : 2.34% “Significantly higher than the Standardization Groups at the p=.05 level. wk kn error of the score the student actually received (the observed score) around 68% of the time. The true score would be within two standard errors of the observed scor approximately 96 of the time. Naturally, the lower the standard error of measure- ment, the more reliable the scores. (chi-squared) tests°* were used ti ‘st the hypotheses that the standard errors of measurement for the test scores in the Virgin Islands sample were greater than those for the Standardization sample. Sixteen of the 108 tests show signifi- cantly higher standard errors for the V.I. sample at the p=.05 level of significance. Nine of them occur in the high school tests in reading and English areas. These tests need to be looked %t closely. Among the remaining seven there seem to be no patterns. I1t should be noted, however, that in two of these cases, Mathematics Applications in grade 4 and Listening Com- prehension in grade 2, the differences are in one district and in the total system scores. Since the total System scores are obviously affected by the individual district scores, it is possible that the large total System standard errors may be a result of the lower reliability obtained from the district scores. x’ /df where df is the number of degrees of freedom, S* is the square of the V.1. sample standard error of measurement, a? is the square of the standardization groups standard error of measurement. (Darlington, 1975) Summary and Conclusions The scores obtained from the testing of a representative sample of U.S. Virgin Islands students using the 1973 edition of the Stanford Achievement Test appear to be both content valid and reliable. This is Significant in that this test, and all standardized tests of academic achievement published in the United States, have been designed without including studies of noncontinental U.S. public school curriculum in the test plan- ning process or using noncontinental U.S. students in its standardization studies. It is clear that the test objectives, as stated by the pub- lisher, are a good match for those used in U.S. Virgin Islands public schools. In addition, the reliability estimates uf the scores obtained from Virgin Islands students are, in most cases, not significantly different from those obtained using the conti- nental U.S. standardization samples. At this point it may be useful to examine the distinction between differences that are “statistically significant" and those that are “educationally Significant."' The statement that two values are "statistically Significant" implies that we are coniident that the difference between the two values is not zero. This is no guarantee that the differences are not trivial. For instance, we may weigh two packages on the same, very accurate, scale and find that one weighs 25 kilograms while the other weighs 25.5 kilograms. If we were trying to decide which of these packages to assign each of two people to carry based on their relative strengths, we could probably conclude that either person could carry J + eS either package. The difference of one half of a kilogram was real (i.e. nonzero), but it was so small that it was trivial. Likewise, differences in reliability estimates noted in this study may be statistically Significant, but so small as to allow us to conclude that the test scores were reliable enough for us to use to make educational decisions (i.e. the differences we-e not educationallv significant). With the exception of the grade 12 and grade 10 Reading test scores and the grade 10 English test scores from the St. Croix district, the differences observed in standard error of measurement estimates seem to be so small as not to be educationally significant. Putting aside the question of the comparability of the obtained reliability estimates between the standardization samples and the Virgin Islands sample, the question of whether or not the scores obtained from the U.S.V.I. sample are reliable enough for us to use them to make educational decisions needs to be addressed. ‘The degree of reliability we demand in our educational measures depends largely on the nature of the decision to be made" (Gronlund, 1976, p.124). Standardized test results are used b [ z school personnel as one source of information for making instruc- tional and curricular decisions. Other sources of information such as teacher made classroom tests and observational techniques are combined with the results of standardized tests before final educational decisions are made in schools. Finally, this partic- ular study was designed to point out strengths and weaknesses in basic skills areas in U.S. Virgin Islands public sciiools. Those a persons entrusted with the responsibility for making curricular and instructional decisions in the Department of Education will information before making changes in what goes on in schools. Further, decisions made will always be open to confirmation and change. Cronbach (1970) points out that the reversability of decisions made on the basis of test data is an important factor to take into consideration in making judgements concerning desired levels of reliability. The reli- ability estimates obtained from the U.S.V.1. sample which seem to cluster from the middle .80's to the middle .90's are more than adequate to allow the confident use of the obtained scores. References Anastasi, A. oaregrencsal Psychology (3rd ed.). New York: Macmillan, 1958. Asher, W.J. neeek sonal Research and Evaluation Methods. Boston: Y id + + 7 R oa a 6: “—_ bei 4 5 A a TR a onl cee i will Le rown « WO. 1976. ae - Review of the Test of Academic Skills In O.K. Buros (Ed.), The E ighth Mental Measurements Yearbo ok (wos. 4}, Highland Park, N.J.: Gryphon Press, 1978. Cronbach, L.J. Essentials of Psychological Testing (3rd ed.). New York: Harper and Row, 1970. Cronbach, 1..J. & Meehl, P.E. Construct validity in a eee tests. ?sychological Bulletin, 1955, 52, 281 Darlington, R.B. Radical and Squares. Ithaca, N.Y.: Logan Hill ‘ Deutsch, M. Minority group and class status as related to social and personality factors in scholastic achievement (Monograph No. 2). Ithaca, N.Y.: The Society for Applied Anthropology, 1960. Ebel, R.L. Must all tests be valid? American Psychologist, 1961, 16, 640-643. Ebel, R.L. Review of the 1973 edition of the Stanford Achievement Test. In O.K. Buros (Ed.), The Eighth Mental Measurements Yearbook (vol.1). Highland Park, N.J.: Gryp hie Press, 1978. Eoei. 2.1. eesent leis of Educational Nessuremeats (3rd eds). Englewood Cliffs, N.J.: Prentice-Hall, French, Jv. & Michael, W.B. (Cochairmen) Standards for Educational and Psychological Tests and Manuals. Washington, D.Cc.: American Psychological Association, 1966. Gronlund, N.E. Measurement and Evaluation in Teaching (3rd ed.). New York: Macmillan, I976. Hayes, W.L. Statistics for the Social Sciences (2nd ed.). New York: Holt, Rinehart, and Winston, 197}. Kasdon, L.M. The Stanford Achievement Test oe of the 1973 edition of the Stanford Achievement Test). Reading Teache 1974, af, F43*, — =’ 4 -46- Lehmann, I.J. The Stanford Achievement Test Series - 1973 (Review of the 1973 edition of the Stanford Achievement Test). Journal of Educational Measurement, 1975, 12, £97-306,, e tee: Passow, A.H. Review of the 1973 edition of the Stanford Achieve- ment Test. In O.K. Buros (Ed.), The Eighth Mental Measure- ments Yearbook (vol. 1). Highland Park, N.J.: Gryphon Press, 1978. Stanford Achievement Test. Technical Data Report. New York: Harcourt Brace Jovanovich, 1975. Thorndike, R.L. Review of the Test of Academic Skills. In O.K. Buros (Ed.), The Eighth Mental Measurements Yearbook (vol. 1). Highland Park, N.J.: Gryphon Press, 1078. Tyler, L. The Psychology of Individual Differences (2nd ed.). New York: Appleton-Century-crofts, 1956. Warwick, D.P. & Lininger, C.A. The Sample Survey: Theory and Practice. New York: McGraw-Hill, 1975.