استفاده از نمونه‌ گیری تصادفی ساده در تحلیل کلاسیک آزمون برای داده‌های‌‌ بزرگ با هدف کاهش زمان پردازش

نوع مقاله : مقاله پژوهشی

نویسنده

گروه آمار، سازمان سنجش آموزش کشور، تهران، ایران

چکیده
هدف این پژوهش یافتن یک روش جدید برای کاهش زمان پردازش داده‌ها و محاسبه شاخص آسانی و شاخص تمیز سؤالات در آزمون‌های بزرگ است. برای این منظور از روش نمونهگیری تصادفی ساده در محاسبه شاخص آسانی و شاخص تمیز در روش تحلیل کلاسیک آزمون و تولید برنامه کامپیوتری برای اجرای نمونه آزمایشی با داده‌های بزرگ استفاده شده است. نتایج نشان داد که استفاده از این روش باعث کاهش زمان برآورد پارامترها می‌شود. برای مثال، سؤالی که با صرف 7393 دقیقه (بیش از 133 ساعت 5 شبانه‌روز) مقدار شاخص P را برابر 50/0 به دست آورده بود، با 10 بار تکرار و صرف زمان متوسط 2/129 دقیقه، مقدار متوسط این شاخص را برابر 503/0 محاسبه کرد. درنتیجه ترکیب روش نمونه‌گیری تصادفی ساده با روش تحلیل کلاسیک آزمون، باعث کاهش چشمگیر زمان پردازش داده‌ها شده و لذا استفاده از این روش به سازمان‌های بزرگ مانند سازمان‌هایی که آزمون‌هایی را در سطح ملی برگزار می‌کنند قویاً توصیه می‌شود.




کلیدواژه‌ها


عنوان مقاله English

Utilizing Simple Random Sampling (SRS) in Classical Test Theory for Large-Scale Assessments: A Time-Efficiency Approach

نویسنده English

Behrooz Kavehie
Department of Statistics, National Organization for Educational Testing, Tehran, Iran
چکیده English

This study proposes a novel method to significantly reduce data processing time for calculating facility (P) and discrimination indices in large-scale assessments. By integrating Simple Random Sampling (SRS) with Classical Test Theory (CTT), we developed a computational approach that minimizes processing time while maintaining statistical accuracy. For illustration, a test question requiring 7,393 minutes (5.5 days) for full analysis yielded a P-index of 0.50. Our SRS-CTT method achieved comparable results (P-index = 0.503) in just 129.2 minutes (mean across 10 iterations), representing a 98.3% reduction in processing time. The accompanying computer program enables efficient analysis of national-level datasets, making this method particularly valuable for high-stakes testing organizations. We strongly recommend its adoption for nationwide educational assessments.

کلیدواژه‌ها English

Classical Test Theory
Large-Scale Dataset
Process Time
Simple Random Sampling
فلسفی‌نژاد، محمدرضا، فرخی، نورعلی، و بهرامی، لیلا (1395). ویژگی‌های روان‌سنجی امتحانات نهایی سال سوم متوسطه و قابلیت آن‌ها در گزینش داوطلبان ورود به دوره‌های کارشناسی. فصلنامه اندازه‌گیری تربیتی، 7(23)، 45-76. https://doi.org/10.22054/jem.2017.6387.1196
یونسی، جلیل، دلاور، علی، و فلسفی‌نژاد، محمدرضا. (1389). بررسی ویژگیهای روان‌سنجی سؤالات تخصصی آزمون‌های فراگیر رشته روان‌شناسی دانشگاه پیام نور در سال 1385. فصلنامه اندازه‌گیری تربیتی، 1 (2)، 139-169. https://jem.atu.ac.ir/article_5636.html
یونسی، جلیل، اسکندری، فرزاد، دلاور، علی، فلسفی‌نژاد، محمدرضا، و فرخی، نورعلی. (1393). مقایسه توانمندی رویکرد بیزی مدل IRT چندسطحی و مدل کلاسیک چندسطحی: تحلیل داده‌های آزمون فیزیک تیمز پیشرفته 2008. فصلنامه اندازه‌گیری تربیتی، 5 (15)، 166-186. https://jem.atu.ac.ir/article_277.html
Adekunle Ayanwale, M., Chere-Masopha, J., & Morena, M. C. (2022). The classical test or item response measurement theory: The status of the framework at the Examination Council of Lesotho. International Journal of Learning, Teaching and Educational Research, 21(8), 409-426. https://doi.org/10.26803/ijlter.21.8.22
Algina, J., & Swaminathan, H. (2015). Psychometrics: Classical test theory. In J. D. Wright (Ed.), International encyclopedia of the social & behavioral sciences (2nd ed., Vol. 19, pp. 423-430). Elsevier. https://doi.org/10.1016/B978-0-08-097086-8.44049-4
Alon, N., Ben-Eliezer, O., Dagan, Y., Moran, S., Naor, M., & Yogev, E. (2021). Adversarial laws of large numbers and optimal regret in online classification. In Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing (pp. 447-455). https://doi.org/10.1145/3406325.3451029
Brown, L. D., Cai, T. T., & DasGupta, A. (2001). Interval estimation for a binomial proportion. Statistical Science, 16(2), 101-133. https://doi.org/10.1214/ss/1009213286
Cochran, W. G. (1977). Sampling techniques (3rd ed.). Wiley. ISBN-13: 978-0471162407
Cunningham, G. K. (1998). Assessment in the classroom: Constructing and interpreting texts. Falmer Press.
Dekking, F. M., Kraaikamp, C., Lopuhaä, H. P., & Meester, L. E. (2005). A modern introduction to probability and statistics. Springer. https://doi.org/10.1007/1-84628-168-7
Domino, G., & Domino, M. L. (2006). Psychological testing: An introduction (2nd ed.). Cambridge University Press. https://doi.org/10.1017/CBO9780511813757
Falsafi-Nejad, M. R., Farrokhi, N., & Bahrami, L. (2016). [Psychometric characteristics of third-grade high school final examinations and their capability in selecting candidates for bachelor's degree programs]. Educational Measurement Quarterly, 7(23), 45-76. https://doi.org/10.22054/jem.2017.6387.1196 [In Persian]
Hazra, A. (2017). Using the confidence interval confidently. Journal of Thoracic Disease, 9(10), 4125-4130. https://doi.org/10.21037/jtd.2017.09.14
Illowsky, B., Dean, S., & Community College Open Educational Resources Group. (2018). Introductory statistics. OpenStax. https://openstax.org/details/books/introductory-statistics
Kelley, T. L. (1939). The selection of upper and lower groups for the validation of test items. Journal of Educational Psychology, 30(1), 17-24. https://doi.org/10.1037/h0057123
Khare, V., Nema, S., & Baredar, P. (2020). Ocean energy modeling and simulation with big data: Computational intelligence for system optimization and grid integration. Elsevier. https://doi.org/10.1016/C2018-0-05113-1
Oosterhof, A. C. (1976). Similarity of various item discrimination indices. Journal of Educational Measurement, 13(2), 145-150. https://doi.org/10.1111/j.1745-3984.1976.tb00005.x
Popham, W. J. (1999). Classroom assessment: What teachers need to know (2nd ed.). Allyn & Bacon.
Rehfisch, J. M. (1958). Some scale and test correlates of a personality rigidity scale. Journal of Consulting Psychology, 22(5), 372-374. https://doi.org/10.1037/h0048746
Wallis, S. A. (2013). Binomial confidence intervals and contingency tests: Mathematical fundamentals and the evaluation of alternative methods. Journal of Quantitative Linguistics, 20(3), 178-208. https://doi.org/10.1080/09296174.2013.799918
Yao, K., & Gao, J. (2016). Law of large numbers for uncertain random variables. IEEE Transactions on Fuzzy Systems, 24(3), 615-621. https://doi.org/10.1109/TFUZZ.2015.2466080
Younesi, J., Delavar, A., & Falsafi-Nejad, M. R. (2010). [Investigating psychometric characteristics of specialized questions in comprehensive exams of psychology at Payame Noor University in 2006]. Educational Measurement Quarterly, 1(2), 139-169. https://jem.atu.ac.ir/article_5636.html [In Persian]
Younesi, J., Eskandari, F., Delavar, A., Falsafi-Nejad, M. R., & Farrokhi, N. (2014). [Comparing the capability of Bayesian approach of multilevel IRT model and classical multilevel model: Analysis of TIMSS advanced 2008 physics test data]. Educational Measurement Quarterly, 5(15), 166-186. https://jem.atu.ac.ir/article_277.html [In Persian]
Zar, J. H. (1999). Biostatistical analysis (4th ed.). Prentice Hall.