Document Type : Research Paper

Authors

1 Professor, Department of Assessment Measurment, Allameh tabataba’i University, Tehran, Iran

2 Education department, pardis bentolhoda Faculty, Farhangiyan University, Tehran, Iran

3 Assistant Professor, Department of Assessment Measurment, Allameh tabataba’i University, Tehran, Iran.

4 Professor, Department of Emergency Medicine, Department of Medical Education, School of Medicine, Tehran University of Medical Sciences, Tehran, Iran

Abstract

The Item Response Theory (IRT) has been extensively used in the development of tools in recent years. The present study aimed at determining the standard and cut- off score of the medical basic sciences comprehensive test using the IRT model, the item-mapping method. The statistical population in this study was all the candidates for the medical basic sciences comprehensive test of the tenth pole of Iran (n=324), and the responses of all candidates were analyzed to determine the standard and cut- off score. The tool used in this research was the basic science comprehensive test with 200 four-option items that was taken in September of 2016. In this cross-sectional study, Win steps was used to analyze the items. The findings of this study showed that the test has an item reliability of PR=0/98 and person reliability of IP=0/92 which indicated that the sample variance and test length were appropriate and the items were appropriately selected from the nine domains. Performance of the reviewers’ panel and evaluation of the items with the item map method showed that the agreement of the reviewers after three review stages included a difficulty parameter of b=1/3 and a raw score of 103. The result of this research showed that the item map method can result in higher agreement among reviewers to determine the cut score.

Keywords

عباسی، هادی. (1393). ارزیابی جامع و تعیین استاندارد علمی چیرگی در آزمون‌های تخصصی ورود به دوره‌های انترنی رشته پزشکی با استفاده از مدل‌های کلاسیک و خصیصه مکنون. رساله دکتری رشته سنجش و اندازه‌گیری دانشگاه علامه طباطبایی.
مرتاض هجری، سارا، جلیلی، محمد، و لباف، علی. (1390). تعیین نمره حدنصاب قبولی آزمون عینی ساختارمند بالینی به روش انگوف و ارزیابی تأثیر بحث و بررسی نمرات واقعی. مجله ایرانی آموزش در علوم پزشکی، 11(8)، 885-894.‎
مینائی، اصغر (1393). استفاده از مدل اندازه‌گیری راش در ارزیابی ویژگی‌های اندازه‌گیری تست مهارت‌های حرکتی (TVMS-R). فصلنامه آموزش‌وپرورش، 5 (18)، 77-114.
وزارت بهداشت، درمان و آموزش پزشکی، معاونت آموزش. (1396). آیین‏نامه آموزشی دوره دکتری عمومی پزشکی مصوب شصت و هفتمین جلسه شورای عالی برنامه‏ریزی علوم پزشکی.
Abbasi, H. (2014). Comprehensive evaluation and determination of scientific mastery standard in specialized entrance exams for medical internship programs using classical and latent trait models [Unpublished doctoral dissertation]. Allameh Tabataba'i University. [In Persian]
Alshawwa, L. (2023). Standard Setting: A Review of Methods. Asian Journal of Education and Social Studies, 42(2), 1-7.
Baron, P., Sireci, S. G., & Slater, S. C. (2021). Evaluating Panelists’ Understanding of Standard Setting Data. Educational Measurement: Issues and Practice, 40(2), 16-25.
Berk, R. A. (1996). Standard setting: The next generation (where few psychometricians have gone before!). Applied measurement in education, 9(3), 215-225.
Boone, W. J., Staver, J. R., & Yale, M. S. (2013). Rasch analysis in the human sciences. Springer Science & Business Media.
Buckendahl, C. W., Smith, R. W., Impara, J. C., & Plake, B. S. (2002). A comparison of Angoff and Bookmark standard setting methods. Journal of Educational measurement, 39(3), 253-263.
Cetin, S., & Gelbal, S. (2013). A Comparison of Bookmark and Angoff Standard Setting Methods. Educational Sciences: Theory and Practice, 13(4), 2169-2175.
Cizek, G. J. (Ed.). (2012). Setting Performance standards: Foundations, methods, and innovations. Routledge.
Duncan, P. W., Bode, R. K., Lai, S. M., Perera, S., & Glycine Antagonist in Neuroprotection Americas Investigators. (2003). Rasch analysis of a new stroke-specific outcome scale: Stroke Impact Scale. Archives of physical medicine and rehabilitation, 84(7), 950-963.
Epstein, R. M., & Hundert, E. M. (2002). Defining and assessing professional competence. Jama, 287(2), 226-235.
Glaser, B. E., Baldwin, P., Margolis, M. J., Mee, J., & Winward, M. (2017). An experimental study of the internal consistency of judgments made in bookmark standard setting. Journal of Educational Measurement, 54(4), 481-497.
Hambleton, R. K., & Pitoniak, M. J. (2006). Setting performance standards. Educational measurement, 4, 433-470.
Hein, S. F., & Skaggs, G. E. (2009). A qualitative investigation of panelists' experiences of standard setting using two variations of the bookmark method. Applied Measurement in Education, 22(3), 207-228.
Jiao, H., Lissitz, R. W., Macready, G., Wang, S., & Liang, S. (2011). Exploring levels of performance using the mixture Rasch model for standard setting1. Psychological Test and Assessment Modeling, 53(4), 499.
Kane, M. T. (1995). Examinee-centered vs. task-centered standard setting. In the Proceedings of the Joint Conference on Standard Setting for Large-scale Assessment (Vol. 2, pp. 119-141).
Kang, Y. (2022). Evaluating the cutoff score of the advanced practice nurse certification examination in Korea. Nurse Education in Practice, 63, 103407.
Linacre, J. M. (2016). Winsteps® Rasch measurement computer program. (No Title).
MacCann, R. G., & Stanley, G. (2019). The use of Rasch modeling to improve standard setting. Practical Assessment, Research, and Evaluation, 11(1), 2.
McKinley, D. W., Newman, L. S., & Wiser, R. F. (1996). Using the Rasch model in the standard setting process. In annual meeting of the National Council of Measurement in Education, New York, NY.
Minaei, A. (2014). Using the Rasch measurement model in evaluating the measurement characteristics of the Test of Gross Motor Skills (TGMS-R). Quarterly Journal of Education, 5(18), 77–114. [In Persian]
Ministry of Health and Medical Education, Deputy of Education. (2017). Educational regulations for the general medicine doctoral program approved by the 67th session of the Supreme Council of Medical Education Planning. [In Persian]
Mitzel, H. C., Lewis, D. M., Patz, R. J., & Green, D. R. (2013). The bookmark procedure: Psychological perspectives. In Setting Performance standards (pp. 263-296). Routledge.
Mortaz Hajri, S., Jalili, M., & Laffaf, A. (2011). Determining the cut-off score for the objective structured clinical examination using the Angoff method and evaluating the impact of real score discussion. Iranian Journal of Medical Education, 11(8), 885–894. [In Persian]
Peterson, C. H., Schulz, E. M., & Engelhard Jr, G. (2011). Reliability and validity of bookmark‐based methods for standard setting: comparisons to Angoff‐based methods in the National Assessment of Educational Progress. Educational Measurement: Issues and Practice, 30(2), 3-14.
Schnabel, S. D. (2018). A comparison of the Angoff and item mapping standard setting methods for a certification examination (Doctoral dissertation, University of Illinois at Chicago).
Shepard, L. A. (1995). Implications for standard setting of the National Academy of Education evaluation of the National Assessment of Educational Progress achievement levels. In Joint conference on standard setting for large-scale assessments. Vol. 2. Proceedings (pp. 143-160).
Tennant, A., & Conaghan, P. G. (2007). The Rasch measurement model in rheumatology: what is it and why use it? When should it be applied, and what should one look for in a Rasch paper? Arthritis Care & Research, 57(8), 1358-1362.
Wagner, P., Hendrich, J., Moseley, G., & Hudson, V. (2007). Defining medical professionalism: a qualitative study. Medical education, 41(3), 288-294.
Wang, N. (2003). Use of the Rasch IRT model in standard setting: An item‐mapping method. Journal of Educational Measurement, 40(3), 231-253.