Utilizing Simple Random Sampling (SRS) in Classical Test Theory for Large-Scale Assessments: A Time-Efficiency Approach

Document Type : Research Paper

Author

Department of Statistics, National Organization for Educational Testing, Tehran, Iran

Abstract
This study proposes a novel method to significantly reduce data processing time for calculating facility (P) and discrimination indices in large-scale assessments. By integrating Simple Random Sampling (SRS) with Classical Test Theory (CTT), we developed a computational approach that minimizes processing time while maintaining statistical accuracy. For illustration, a test question requiring 7,393 minutes (5.5 days) for full analysis yielded a P-index of 0.50. Our SRS-CTT method achieved comparable results (P-index = 0.503) in just 129.2 minutes (mean across 10 iterations), representing a 98.3% reduction in processing time. The accompanying computer program enables efficient analysis of national-level datasets, making this method particularly valuable for high-stakes testing organizations. We strongly recommend its adoption for nationwide educational assessments.

Keywords


فلسفی‌نژاد، محمدرضا، فرخی، نورعلی، و بهرامی، لیلا (1395). ویژگی‌های روان‌سنجی امتحانات نهایی سال سوم متوسطه و قابلیت آن‌ها در گزینش داوطلبان ورود به دوره‌های کارشناسی. فصلنامه اندازه‌گیری تربیتی، 7(23)، 45-76. https://doi.org/10.22054/jem.2017.6387.1196
یونسی، جلیل، دلاور، علی، و فلسفی‌نژاد، محمدرضا. (1389). بررسی ویژگیهای روان‌سنجی سؤالات تخصصی آزمون‌های فراگیر رشته روان‌شناسی دانشگاه پیام نور در سال 1385. فصلنامه اندازه‌گیری تربیتی، 1 (2)، 139-169. https://jem.atu.ac.ir/article_5636.html
یونسی، جلیل، اسکندری، فرزاد، دلاور، علی، فلسفی‌نژاد، محمدرضا، و فرخی، نورعلی. (1393). مقایسه توانمندی رویکرد بیزی مدل IRT چندسطحی و مدل کلاسیک چندسطحی: تحلیل داده‌های آزمون فیزیک تیمز پیشرفته 2008. فصلنامه اندازه‌گیری تربیتی، 5 (15)، 166-186. https://jem.atu.ac.ir/article_277.html
Adekunle Ayanwale, M., Chere-Masopha, J., & Morena, M. C. (2022). The classical test or item response measurement theory: The status of the framework at the Examination Council of Lesotho. International Journal of Learning, Teaching and Educational Research, 21(8), 409-426. https://doi.org/10.26803/ijlter.21.8.22
Algina, J., & Swaminathan, H. (2015). Psychometrics: Classical test theory. In J. D. Wright (Ed.), International encyclopedia of the social & behavioral sciences (2nd ed., Vol. 19, pp. 423-430). Elsevier. https://doi.org/10.1016/B978-0-08-097086-8.44049-4
Alon, N., Ben-Eliezer, O., Dagan, Y., Moran, S., Naor, M., & Yogev, E. (2021). Adversarial laws of large numbers and optimal regret in online classification. In Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing (pp. 447-455). https://doi.org/10.1145/3406325.3451029
Brown, L. D., Cai, T. T., & DasGupta, A. (2001). Interval estimation for a binomial proportion. Statistical Science, 16(2), 101-133. https://doi.org/10.1214/ss/1009213286
Cochran, W. G. (1977). Sampling techniques (3rd ed.). Wiley. ISBN-13: 978-0471162407
Cunningham, G. K. (1998). Assessment in the classroom: Constructing and interpreting texts. Falmer Press.
Dekking, F. M., Kraaikamp, C., Lopuhaä, H. P., & Meester, L. E. (2005). A modern introduction to probability and statistics. Springer. https://doi.org/10.1007/1-84628-168-7
Domino, G., & Domino, M. L. (2006). Psychological testing: An introduction (2nd ed.). Cambridge University Press. https://doi.org/10.1017/CBO9780511813757
Falsafi-Nejad, M. R., Farrokhi, N., & Bahrami, L. (2016). [Psychometric characteristics of third-grade high school final examinations and their capability in selecting candidates for bachelor's degree programs]. Educational Measurement Quarterly, 7(23), 45-76. https://doi.org/10.22054/jem.2017.6387.1196 [In Persian]
Hazra, A. (2017). Using the confidence interval confidently. Journal of Thoracic Disease, 9(10), 4125-4130. https://doi.org/10.21037/jtd.2017.09.14
Illowsky, B., Dean, S., & Community College Open Educational Resources Group. (2018). Introductory statistics. OpenStax. https://openstax.org/details/books/introductory-statistics
Kelley, T. L. (1939). The selection of upper and lower groups for the validation of test items. Journal of Educational Psychology, 30(1), 17-24. https://doi.org/10.1037/h0057123
Khare, V., Nema, S., & Baredar, P. (2020). Ocean energy modeling and simulation with big data: Computational intelligence for system optimization and grid integration. Elsevier. https://doi.org/10.1016/C2018-0-05113-1
Oosterhof, A. C. (1976). Similarity of various item discrimination indices. Journal of Educational Measurement, 13(2), 145-150. https://doi.org/10.1111/j.1745-3984.1976.tb00005.x
Popham, W. J. (1999). Classroom assessment: What teachers need to know (2nd ed.). Allyn & Bacon.
Rehfisch, J. M. (1958). Some scale and test correlates of a personality rigidity scale. Journal of Consulting Psychology, 22(5), 372-374. https://doi.org/10.1037/h0048746
Wallis, S. A. (2013). Binomial confidence intervals and contingency tests: Mathematical fundamentals and the evaluation of alternative methods. Journal of Quantitative Linguistics, 20(3), 178-208. https://doi.org/10.1080/09296174.2013.799918
Yao, K., & Gao, J. (2016). Law of large numbers for uncertain random variables. IEEE Transactions on Fuzzy Systems, 24(3), 615-621. https://doi.org/10.1109/TFUZZ.2015.2466080
Younesi, J., Delavar, A., & Falsafi-Nejad, M. R. (2010). [Investigating psychometric characteristics of specialized questions in comprehensive exams of psychology at Payame Noor University in 2006]. Educational Measurement Quarterly, 1(2), 139-169. https://jem.atu.ac.ir/article_5636.html [In Persian]
Younesi, J., Eskandari, F., Delavar, A., Falsafi-Nejad, M. R., & Farrokhi, N. (2014). [Comparing the capability of Bayesian approach of multilevel IRT model and classical multilevel model: Analysis of TIMSS advanced 2008 physics test data]. Educational Measurement Quarterly, 5(15), 166-186. https://jem.atu.ac.ir/article_277.html [In Persian]
Zar, J. H. (1999). Biostatistical analysis (4th ed.). Prentice Hall.