Antithesis of Human Rater: Psychometric Responding to Shifts Competency Test Assessment Using Automation (AES System)

  • Mohammad Idhom Doctoral Program of Vocational Education, Universitas Negeri Surabaya, Indonesia
  • I Gusti Putu Asto Buditjahjanto Doctoral Program of Vocational Education, Universitas Negeri Surabaya, Indonesia
  • Munoto Doctoral Program of Vocational Education, Universitas Negeri Surabaya, Indonesia
  • Trimono Data Science Study Program, UPN "Veteran" Jawa Timur, Indonesia
  • Prismahardi Aji Riyantoko Data Science Study Program, UPN "Veteran" Jawa Timur, Indonesia
Keywords: Automated essay scoring, Competency, Human rater, Information and communication technology, Vocational education

Abstract

This research is part of proof tests to a combination of statistical processing methods, collecting assessment rubrics in vocational education by comparing two systems, automated essay scoring and human rater. It aims to analyze the final assessment score of essays in Akademi Komunitas Negeri (AKN) Pacitan (Pacitan’s State Community College) and Akademi Komunitas Negeri (AKN) Blitar (Blitar’s State Community College) in East Java, Indonesia. The provisional assumption is that the results show an antithesis to the assessment of human feedback with an automated system due to the conversion of scores between the rubric and the algorithm design. As the hypothesis, algorithm-based score conversion affects automated essay scoring and human rater methods, which led to antithesis feedback. The validity and reliability of the measurement maintain the scoring consistency between the two methods and the accuracy of the answers. The novelty of this article is comparing between AES system and Human Rater using statistical methods. The research shows that there is a similar result using the psychometrics approach, which indicates different metaphor expressions and language systems. Thus, the objective of this study is to provide assistance in the advancement of an information technology system that utilizes a scoring mechanism merging computer and human evaluations, employing a psychological approach known as psychometric leads.

Downloads

Download data is not yet available.

References

Abidin, S. N. Z., & Jaffar, M. M. (2014). Forecasting share prices of small size companies in bursa Malaysia using geometric Brownian motion. Applied Mathematics & Information Sciences, 8(1), 107–112. https://doi.org/10.12785/amis/080112

Almeida, F., & Buzady, Z. (2023). exploring the impact of a serious game in the academic success of entrepreneurship students. Journal of Educational Technology Systems, 51(4), 436–454. https://doi.org/10.1177/00472395231153187

Atteveldt, W. v., Velden, M. A. C. G. v. D., & Boukes, M. (2021). The validity of sentiment analysis: Comparing manual annotation, crowd-coding, dictionary approaches, and machine learning algorithms. Communication Methods and Measures, 15(2), 121–140. https://doi.org/10.1080/19312458.2020.1869198

Buditjahjanto, I. G. P. A., Idhom, M., Munoto, M., & Samani, M. (2022). An automated essay scoring based on neural networks to predict and classify competence of examinees in community academy. TEM Journal, 11(4), 1694–1701. https://doi.org/10.18421/TEM114-34

Burrows, S., Gurevych, I., & Stein, B. (2015). The eras and trends of automatic short answer grading. International Journal of Artificial Intelligence in Education, 25(1), 60–117. https://doi.org/10.1007/s40593-014-0026-8

Cassidy, D. T. (2016). A multivariate student’s t-distribution. Open Journal of Statistics, 6(3), 443–450. https://doi.org/10.4236/ojs.2016.63040

Collier-Sewell, F., Atherton, I., Mahoney, C., Kyle, R. G., Hughes, E., & Lasater, K. (2023). Competencies and standards in nurse education: The irresolvable tensions. Nurse Education Today, 125, 105782. https://doi.org/10.1016/j.nedt.2023.105782

Dong, F., Zhang, Y., & Yang, J. (2017). Attention-based recurrent convolutional neural network for automatic essay scoring. Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), 153–162. https://doi.org/10.18653/v1/K17-1017

Facione, P. A. (2015). Critical thinking: What it is and why it counts. Insight Assessement. https://www.insightassessment.com/wp-content/uploads/ia/pdf/whatwhy.pdf

Ghorbani, H. (2019). Mahalanobis distance and its application for detecting multivariate outliers. Facta Universitatis Series: Mathematics and Informatics, 34(3) 583-595. https://doi.org/10.22190/FUMI1903583G

Grover, G., Sabharwal, A., & Mittal, J. (2014). Application of multivariate and bivariate normal distributions to estimate duration of diabetes. International Journal of Statistics and Applications, 4(1), 46–57.

Haley, B., Heo, S., Wright, P., Barone, C., Rao Rettiganti, M., & Anders, M. (2017). Relationships among active listening, self-awareness, empathy, and patient-centered care in associate and baccalaureate degree nursing students. NursingPlus Open, 3(2017), 11–16. https://doi.org/10.1016/j.npls.2017.05.001

Hasanah, U., Permanasari, A. E., Kusumawardani, S. S., & Pribadi, F. S. (2019). A scoring rubric for automatic short answer grading system. Telkomnika (Telecommunication Computing Electronics and Control), 17(2), 763-770. https://doi.org/10.12928/telkomnika.v17i2.11785

Heale, R., & Twycross, A. (2015). Validity and reliability in quantitative studies. Evidence Based Nursing, 18(3), 66–67. https://doi.org/10.1136/eb-2015-102129

Kennedy, I. (2022). Sample size determination in test-retest and cronbach alpha reliability estimates. British Journal of Contemporary Education, 2(1), 17–29.

Liang, Y., Coelho, C. A., & von Rosen, T. (2022). Hypothesis testing in multivariate normal models with block circular covariance structures. Biometrical Journal, 64(3), 557–576. https://doi.org/10.1002/bimj.202100023

Mao, L., Liu, O. L., Roohr, K., Belur, V., Mulholland, M., Lee, H.-S., & Pallant, A. (2018). Validation of automated scoring for a formative assessment that employs scientific argumentation. Educational Assessment, 23(2), 121–138. https://doi.org/10.1080/10627197.2018.1427570

Meeker, W. Q., Escobar, L. A., & Pascual, F. G. (2022). Statistical methods for reliability data. John Wiley & Sons.

Moses, R. N., & Yamat, H. (2021). Testing the validity and reliability of a writing skill assessment. International Journal of Academic Research in Business and Social Sciences, 11(4). https://doi.org/10.6007/IJARBSS/v11-i4/9028

Nabillah, I., & Ranggadara, I. (2020). Mean absolute percentage error untuk evaluasi hasil prediksi komoditas laut. JOINS (Journal of Information System), 5(2), 250–255. https://doi.org/10.33633/joins.v5i2.3900

Nanni, A. C., & Wilkinson, P. J. (2015). Assessment of ELLs’ critical thinking using the holistic critical thinking scoring rubric. Language Education in Asia, 5(2), 283–291. https://doi.org/10.5746/LEiA/14/V5/I2/A09/Nanni_Wilkinson

Navarro, G., Baeza-Yates, R., & Ribeiro-Neto, B. (1999). Modern information retrieval. ACM Press.

Puñal, O., Aktaş, I., Schnelke, C. J., Abidin, G., Wehrle, K., & Gross, J. (2014). Machine learning-based jamming detection for IEEE 802.11: Design and experimental evaluation. Proceeding of IEEE International Symposium on a World of Wireless, Mobile and Multimedia Networks 2014, 1–10. https://doi.org/10.1109/WoWMoM.2014.6918964

Purnama, I. A. (2015). Pengaruh skema kompensasi denda terhadap kinerja dengan risk preference sebagai variabel moderating. Jurnal Nominal, 4(1), 129–145.

Safiullin, R., Marusin, A., Safiullin, R., & Ablyazov, T. (2019). Methodical approaches for creation of intelligent management information systems by means of energy resources of technical facilities. E3S Web of Conferences, 140, 10008. https://doi.org/10.1051/e3sconf/201914010008

Shekar, B. H., & Dagnew, G. (2019). Grid search-based hyperparameter tuning and classification of microarray cancer data. 2019 Second International Conference on Advanced Computational and Communication Paradigms (ICACCP), 1–8. https://doi.org/10.1109/ICACCP.2019.8882943

Surucu, L., & Maslakci, A. (2020). Validity and reliability in quantitative research. Business & Management Studies: An International Journal, 8(3), 2694–2726. https://doi.org/10.15295/bmij.v8i3.1540

Tong, D. H., Uyen, B. P., & Quoc, N. V. A. (2021). The improvement of 10th students’ mathematical communication skills through learning ellipse topics. Heliyon, 7(11), e08282. https://doi.org/10.1016/j.heliyon.2021.e08282

Wang, Z. (2012). Investigation of the effects of scoring designs and rater severity on students’ ability estimation using different rater models. Conference: 2012 Annual Meeting of the National Council on Measurement in Education.

Wang, Z., Zechner, K., & Sun, Y. (2018). Monitoring the performance of human and automated scores for spoken responses. Language Testing, 35(1), 101-120.

Watkins, S. C. (2020). Simulation-based training for assessment of competency, certification, and maintenance of certification. In J. T. Paige, S. C. Sonesh, D. D. Garbee, L. S. Bonanno (Eds.), Comprehensive Healthcare Simulation: InterProfessional Team Training and Simulation (pp. 225–245). Springer. https://doi.org/10.1007/978-3-030-28845-7_15

Williamson, D. M., Xi, X., & Breyer, F. J. (2012). A framework for evaluation and use of automated scoring. Educational Measurement: Issues and Practice, 31(1), 2–13. https://doi.org/10.1111/j.1745-3992.2011.00223.x

Wong, W. S., & Bong, C. H. (2019). A study for the development of automated essay scoring (AES) in Malaysian English test environment. International Journal of Innovative Computing, 9(1). https://doi.org/10.11113/ijic.v9n1.220

Yoo, K., Rosenberg, M. D., Noble, S., Scheinost, D., Constable, R. T., & Chun, M. M. (2019). Multivariate approaches improve the reliability and validity of functional connectivity and prediction of individual behaviors. NeuroImage, 197, 212–223. https://doi.org/10.1016/j.neuroimage.2019.04.060

Zhang, M. (2013). Contrasting automated and human scoring of essays. R & D Connections, 21(2), 1–11.

Zhou, H., Deng, Z., Xia, Y., & Fu, M. (2016). A new sampling method in particle filter based on pearson correlation coefficient. Neurocomputing, 216, 208–215. https://doi.org/10.1016/j.neucom.2016.07.036

Published
2023-08-31
How to Cite
Idhom, M., Buditjahjanto, I. G. P. A., Munoto, Trimono, & Riyantoko, P. A. (2023). Antithesis of Human Rater: Psychometric Responding to Shifts Competency Test Assessment Using Automation (AES System). Studies in Learning and Teaching, 4(2), 329-340. https://doi.org/10.46627/silet.v4i2.291
Section
Articles
Abstract viewed = 663 times
PDF downloaded = 554 times