Current - Issue
Original Article
AI-Based Question Paper Quality Assessment
Harshal B. Patil1
Shubham M. Koshti2
Rita P. Kurkure3
Dr. Yogesh N. Chaudhari4
Dr. Dhanpal N. Waghulde5
1 2 3 4 Assistant Professor, KCES’s Institute of Management and Research, Jalgaon, Maharashtra, India. 5 Associate Professor, KCES’s Institute of Management and Research, Jalgaon, Maharashtra, India.
Published Online: May-June 2026
Pages: 216-221
Cite this article
↗ https://www.doi.org/10.59256/ijsreat.20260603032References
1. Anderson, L. W., & Krathwohl, D. R. (2001). A taxonomy for learning, teaching, and assessing: A revision of Bloom's educational
objectives. Longman.
2. Alsubait, T., Parsia, B., & Sattler, U. (2016). Measuring similarity in ontologies: A new family of measures. International Journal of
Artificial Intelligence in Education, 26(1), 59-100.
3. Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater v.2. Journal of Technology, Learning, and Assessment, 4(3).
4. Benedetto, L., Crammer, K., & Donat, A. (2021). R2DE: A NLP approach to estimating IRT parameters of newly generated questions.
Proceedings of the 11th International Conference on Learning Analytics and Knowledge (LAK'21), 361-370.
5. Denny, P., Luxton-Reilly, A., & Tempero, E. (2008). All syntax errors are not equal: Towards a classification of novice programming
errors. Proceedings of the 13th Annual Conference on Innovation and Technology in Computer Science Education (ITiCSE), 346-350.
6. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language
understanding. Proceedings of NAACL-HLT 2019, 4171-4186.
7. Foltz, P. W., Laham, D., & Landauer, T. K. (1999). Automated essay scoring: Applications to educational technology. Proceedings of
World Conference on Educational Multimedia, Hypermedia and Telecommunications, 939-944.
8. Gierl, M. J., Lai, H., Hogan, J., & Matovinovic, D. (2017). A method for generating educational test items that are aligned to the common
core state standards. Journal of Applied Testing Technology, 15(1), 1-18.
9. Heilman, M., & Smith, N. A. (2010). Good question! Statistical ranking for question generation. Proceedings of NAACL-HLT 2010, 609-
617.
10. Huang, Z., Liu, Q., Chen, E., Zhao, H., Gao, M., Wei, S., Su, Y., & Hu, G. (2017). Question difficulty prediction for reading problems in
standard tests. Proceedings of the 31st AAAI Conference on Artificial Intelligence (AAAI-17), 1352-1359.
11. Lalor, J. P., Wu, H., & Yu, H. (2019). Learning latent parameters without human response patterns: Item response theory with artificial
crowds. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP), 4240-4250.
12. Li, X., Zhang, H., & Zhou, Y. (2022). Topic-modelling-based syllabus coverage analysis for automated examination evaluation. Computers
& Education, 178, 104405.13. Mitkov, R., & Ha, L. A. (2003). Computer-aided generation of multiple-choice tests. Proceedings of the HLT-NAACL 2003 Workshop on
Building Educational Applications Using NLP, 17-22.
14. Noy, N. F., & McGuinness, D. L. (2001). Ontology development 101: A guide to creating your first ontology. Stanford Knowledge Systems
Laboratory Technical Report KSL-01-05.
15. Page, E. B. (1966). The imminence of grading essays by computer. Phi Delta Kappan, 47(5), 238-243.
16. Peng, J., Wang, L., & Zhao, Y. (2023). Holistic examination quality assessment using multi-dimensional AI evaluation: A large-scale
empirical study. Expert Systems with Applications, 214, 119151.
17. Rodriguez, C., Gutierrez, F., & Deco, C. (2021). Knowledge graph-based curriculum alignment and examination coverage evaluation. IEEE
Transactions on Learning Technologies, 14(4), 501-514.
18. Settles, B., LaFlair, G. T., & Hagiwara, M. (2020). Machine learning–driven language assessment. Transactions of the Association for
Computational Linguistics, 8, 247-263.
19. Tsangaratos, P., Ilia, I., & Loupasakis, C. (2020). Automated mapping of examination learning outcomes to Bloom's Taxonomy using
transfer learning. Applied Sciences, 10(21), 7571.
20. Yahya, A. A., & Osman, A. (2012). Automatic classification of questions in Bloom's taxonomy based on question structure. International
Journal of Engineering Research and Technology, 1(3), 1-6.
objectives. Longman.
2. Alsubait, T., Parsia, B., & Sattler, U. (2016). Measuring similarity in ontologies: A new family of measures. International Journal of
Artificial Intelligence in Education, 26(1), 59-100.
3. Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater v.2. Journal of Technology, Learning, and Assessment, 4(3).
4. Benedetto, L., Crammer, K., & Donat, A. (2021). R2DE: A NLP approach to estimating IRT parameters of newly generated questions.
Proceedings of the 11th International Conference on Learning Analytics and Knowledge (LAK'21), 361-370.
5. Denny, P., Luxton-Reilly, A., & Tempero, E. (2008). All syntax errors are not equal: Towards a classification of novice programming
errors. Proceedings of the 13th Annual Conference on Innovation and Technology in Computer Science Education (ITiCSE), 346-350.
6. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language
understanding. Proceedings of NAACL-HLT 2019, 4171-4186.
7. Foltz, P. W., Laham, D., & Landauer, T. K. (1999). Automated essay scoring: Applications to educational technology. Proceedings of
World Conference on Educational Multimedia, Hypermedia and Telecommunications, 939-944.
8. Gierl, M. J., Lai, H., Hogan, J., & Matovinovic, D. (2017). A method for generating educational test items that are aligned to the common
core state standards. Journal of Applied Testing Technology, 15(1), 1-18.
9. Heilman, M., & Smith, N. A. (2010). Good question! Statistical ranking for question generation. Proceedings of NAACL-HLT 2010, 609-
617.
10. Huang, Z., Liu, Q., Chen, E., Zhao, H., Gao, M., Wei, S., Su, Y., & Hu, G. (2017). Question difficulty prediction for reading problems in
standard tests. Proceedings of the 31st AAAI Conference on Artificial Intelligence (AAAI-17), 1352-1359.
11. Lalor, J. P., Wu, H., & Yu, H. (2019). Learning latent parameters without human response patterns: Item response theory with artificial
crowds. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP), 4240-4250.
12. Li, X., Zhang, H., & Zhou, Y. (2022). Topic-modelling-based syllabus coverage analysis for automated examination evaluation. Computers
& Education, 178, 104405.13. Mitkov, R., & Ha, L. A. (2003). Computer-aided generation of multiple-choice tests. Proceedings of the HLT-NAACL 2003 Workshop on
Building Educational Applications Using NLP, 17-22.
14. Noy, N. F., & McGuinness, D. L. (2001). Ontology development 101: A guide to creating your first ontology. Stanford Knowledge Systems
Laboratory Technical Report KSL-01-05.
15. Page, E. B. (1966). The imminence of grading essays by computer. Phi Delta Kappan, 47(5), 238-243.
16. Peng, J., Wang, L., & Zhao, Y. (2023). Holistic examination quality assessment using multi-dimensional AI evaluation: A large-scale
empirical study. Expert Systems with Applications, 214, 119151.
17. Rodriguez, C., Gutierrez, F., & Deco, C. (2021). Knowledge graph-based curriculum alignment and examination coverage evaluation. IEEE
Transactions on Learning Technologies, 14(4), 501-514.
18. Settles, B., LaFlair, G. T., & Hagiwara, M. (2020). Machine learning–driven language assessment. Transactions of the Association for
Computational Linguistics, 8, 247-263.
19. Tsangaratos, P., Ilia, I., & Loupasakis, C. (2020). Automated mapping of examination learning outcomes to Bloom's Taxonomy using
transfer learning. Applied Sciences, 10(21), 7571.
20. Yahya, A. A., & Osman, A. (2012). Automatic classification of questions in Bloom's taxonomy based on question structure. International
Journal of Engineering Research and Technology, 1(3), 1-6.
Related Articles
2026
Fake Currency Detection Using Deep Learning
2026
Smart E-Commerce System with Dynamic Pricing
2026
Personal Expense Tracker with Currency Converter
2026
Paw Safe: An Extensive Technology-Driven Framework for Stray Dog Rescue, Healthcare Management, Community Engagement, and Smart Urban Governance
2026
Design and Development of a Full-Stack E-Commerce Website
2026