Research & Expertise

Research & Expertise

This page presents research supported by Measurement Incorporated as well as unaffiliated research conducted by current MI staff—illustrating our staff's wide-ranging expertise in educational measurement.

Open Access

Featured Research

Industry-leading, open-access publications authored by our staff.

Huang, Y., Palermo, C., & Wilson, J. (2026). Accuracy and fairness of generative AI in automated essay scoring: Comparing GPT-4o, feature-based models, and human raters. Assessing Writing, 69. https://doi.org/10.1016/j.asw.2026.101047

Abstract. This study examines the accuracy and fairness of generative AI–based automated essay scoring (AES) in a developmentally emergent writing context, compared with a feature-based scoring engine and human ratings. Despite growing interest in generative AI for AES, limited research has examined model performance in early-grade writing, where linguistic and transcription skills are still developing. Using 1768 essays from Grades 3–4, we evaluated GPT-4o under three prompting strategies and a fine-tuned configuration. Human–machine agreement was assessed against human–human benchmarks at both holistic and trait levels using quadratic weighted kappa, exact and adjacent agreement rates, and classification metrics. Fairness was examined using overall score accuracy, overall score difference, and conditional score difference. Results indicate that carefully designed prompting improves GPT-4o's scoring accuracy, approaching human and feature-based AES performance, while the fine-tuned model achieved the closest alignment to human–human agreement. Trait-level analyses revealed complementary strengths: the best-performing prompting approach aligned more closely on higher-level traits (e.g., organization and style), whereas the feature-based engine showed stronger alignment on development of ideas and all surface-level traits. Fairness analyses indicated that no model was entirely free of subgroup differences. These findings suggest that generative AI–based AES can approximate human scoring in early-grade writing when appropriately calibrated.

Huang, Y., Palermo, C., Liu, R., & He, Y. (2025). An early review of generative language models in automated writing evaluation: Advancements, challenges, and future directions for automated essay scoring and feedback generation. Chinese/English Journal of Educational Measurement and Evaluation, 6(2), Article 5. https://doi.org/10.59863/FAMJ7696

Abstract. Automated writing evaluation (AWE) has long supported assessment and instruction, yet existing systems struggle to capture deeper rhetorical and pedagogical aspects of student writing. Recent advances in generative language models (GLMs) such as GPT and Llama present new opportunities, but their effectiveness remains uncertain. This review synthesizes 29 studies on automated essay scoring and 14 on automated writing feedback generation, examining how GLMs are applied through prompting, fine-tuning, and adaptation. Findings show GLMs can approximate human scoring and deliver richer, rubric-aligned feedback, but fairness, validity, and ethical issues remain largely unaddressed. We conclude that GLMs hold promise to enhance AWE, provided that future work establishes robust evaluation frameworks and safeguards to ensure responsible, equitable use.

Insights

White Papers

Our collection of white papers, written by our own industry experts and psychometricians, includes the latest industry research and best practices.

The Quest for Consistency: Double-Scoring Policies and Impacts on Fairness

"Double-scoring a proportion of responses serves the goal of measuring scoring consistency, but when the process includes third readings or resolutions score comparability is undermined. Alternatives that ensure scoring quality and fairness should be a priority in assessment programs."

Download PDF (189.12 KB)

A Gentle Introduction to Automated Scoring

"Automated scoring also offers a variety of benefits for assessment of learning. One benefit is that it is much faster than scoring by teachers or professional raters; once models have been generated, responses can be scored in seconds. This allows assessment results to be available to stakeholders very rapidly. A second benefit is that automated scoring tends to be as accurate or more accurate than multiple professional raters. Furthermore, automated-scoring engines are perfectly reliable in ways that raters are not—an automated-scoring engine will assign the same score to a response every time."

Download PDF (222.99 KB)

White Paper: PEG Changes

"MI continues to monitor advancements in the automated essay scoring field while searching for ways to make PEG as effective as possible in helping students learn to write. As a result, PEG will be ever-evolving."

Download PDF (49.75 KB)

The Case for Professional Learning Communities

"A Professional Learning Community (PLC) is a small group of professionals who continuously seek cutting-edge ideas and collaboratively evaluate how to best apply the new information to the work. The PLC operates under the assumption that to stay ahead of the competition, an organization must learn faster than the competition and consistently produce exceptional work."

Download PDF (883.9 KB)

The Future of Testing

"The future still looks a lot like it did 25 years ago: cognitive-based assessment, online assessment, widespread use of computer adaptive testing, universal access to technology, and instantaneous reporting of test results. So many wonderful things, still within our view but just beyond our grasp!"

Download PDF (486.36 KB)

It Takes Three

"Making sure all students are college and career ready requires not only an alignment of curriculum and instruction with college and career requirements but also an approach to monitoring student progress on a continual basis, with in-class formative assessments, frequent interim assessments, and focused summative assessments. Taken together, formative, interim, and summative assessments, aligned to Common Core State Standards (CCSS), will support instructional decision making and enhance daily learning activities."

Download PDF (315.75 KB)

Aligning Curriculum, Assessment, and Instruction

"A key component of educational achievement test validation is alignment of the test to both curriculum and instruction. By alignment, we mean the degree to which the items of the test, both individually and collectively, match the structure and intent of the curriculum and instruction."

Download PDF (474.81 KB)
Peer-Reviewed

Publications

Peer-reviewed scholarly works by our staff.

2026

Huang, Y., Palermo, C., & Wilson, J. (2026). Accuracy and fairness of generative AI in automated essay scoring: Comparing GPT-4o, feature-based models, and human raters. Assessing Writing, 69, 101047. https://doi.org/10.1016/j.asw.2026.101047

Palermo, C., He, Y., Justice, D., & Vaughn, D. (under review). Improving construct representation in automated scoring through synthetic response generation and validation (invited paper). Journal of Educational Measurement.

Yan, D., Palermo, C., & He, Y. (under review). Automated essay scoring in the age of AI (invited chapter). In A. von Davier & D. Yan (Eds.), Artificial intelligence applications in educational learning and assessment (forthcoming). Springer.

2025

Gui, Y. (2025). Develop a generic essay scorer for practice writing tests of statewide assessments. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con) (pp. 58–81). https://aclanthology.org/2025.aimecon-main.8/

Gui, Y. (2025). From entropy to generalizability: Strengthening automated essay scoring reliability and sustainability. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con) (pp. 312–328). https://aclanthology.org/2025.aimecon-main.34/

Huang, Y., Palermo, C., & Wilson, J. (2025). Identifying active ingredients and uptake patterns in the implementation of an AI-based writing support tool: Insights from a randomized controlled trial. Computers and Education: Artificial Intelligence, 9, 100479. https://doi.org/10.1016/j.caeai.2025.100479

Huang, Y., Palermo, C., Liu, R., & He, Y. (2025). An early review of generative language models in automated writing evaluation: Advancements, challenges, and future directions for automated essay scoring and feedback generation. Chinese/English Journal of Educational Measurement and Evaluation, 6(2), Article 5. https://doi.org/10.59863/FAMJ7696

Huang, Y. & Wilson, J. (2025). Exploring the effectiveness of large-scale automated writing evaluation implementation on state test performance using generalized boosted modeling. Journal of Computer Assisted Learning, 41, e70009. https://doi.org/10.1111/jcal.70009

Huang, Y. & Wilson, J. (2025). Evaluating LLM-based automated essay scoring: Accuracy, fairness, and validity. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress (pp. 71–83). National Council on Measurement in Education. Pittsburgh, PA. https://aclanthology.org/2025.aimeconwip.9/

Huang, Y., Wilson, J., & May, H. (2025). Exploring the long-term effects of the statewide implementation of an automated writing evaluation system on students' state test ELA performance. International Journal of Artificial Intelligence in Education, 35, 1528–1559. https://doi.org/10.1007/s40593-024-00443-9

Jiang, N., Huang, Y., & Chen, J. (2025). Comparison of AI and human scoring on a visual arts assessment. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress (pp. 147–154). National Council on Measurement in Education. Pittsburgh, PA. https://aclanthology.org/2025.aimeconwip.18/

Palermo, C., Chen, T., & Wibowo, A. (2025). Operational alignment of confidence-based flagging methods in automated scoring. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers (pp. 56–60). National Council on Measurement in Education. https://aclanthology.org/2025.aimecon-sessions.6/

2024

Tomkowicz, J., Porter, A., & Palermo, C. (2024). Beyond the clock: Testing time and performance realities. Journal of Applied Testing Technology, 25(1), 26–45. https://www.jattjournal.net/index.php/atp/article/view/173216

Palermo, C., & Wibowo, A. (2024). Automated essay evaluation at scale: Hybrid automated scoring/hand scoring in the summative assessment program. In M. Shermis & J. Wilson (Eds.), The Routledge International Handbook of Automated Essay Evaluation (pp. 23–39). New York, NY: Routledge. https://doi.org/10.4324/9781003397618-3

Wilson, J., & Huang, Y. (2024). Validity of automated essay scores for elementary-age English language learners: Evidence of bias? Assessing Writing, 60, 100815. https://doi.org/10.1016/j.asw.2024.100815

2023

Cui, Z. & He, Y. (2023). Practical considerations in choosing an anchor test form for equating under the random groups design. Measurement: Interdisciplinary Research and Perspectives, 2, 101–113. https://doi.org/10.1080/15366367.2022.2087354

2022

Palermo, C. (2022). Rater characteristics, response content, and scoring contexts: Decomposing determinates of scoring accuracy. Frontiers in Psychology, 13:937097. https://doi.org/10.3389/fpsyg.2022.937097

2021

Clauser, B. E., & Bunch, M. B. (Eds.). (2021). The history of educational measurement: Key advancements in theory, policy, and practice. Routledge. https://doi.org/10.4324/9780367815318

Huang, Y., & Wilson, J. (2021). Using automated feedback to develop writing proficiency. Computers and Composition, 62, 102675. https://doi.org/10.1016/j.compcom.2021.102675

2020

He, Y. & Cui, Z. (2020). Evaluating robust scale transformation methods with multiple outlying common items under IRT true score equating. Applied Psychological Measurement, 44, 296–310. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7262993/

Wang, W., Chen, J. & Kingston, N. (2020). How well do simulation studies inform decisions about multistage testing? Journal of Applied Measurement, 21(3), 271–281. PMID: 33983899.

2019

Murray, A. K., Daoust, C. J., & Chen, J. (2019). Developing instruments to measure Montessori instructional practices. Journal of Montessori Research, 5(1), 50–87. https://doi.org/10.17161/jomr.v5i1.9797

Palermo, C., Bunch, M., & Ridge, K. (2019). Scoring stability in a large-scale assessment program: A longitudinal analysis of leniency/severity effects. Journal of Educational Measurement, 56(3), 626–652. https://doi.org/10.1111/jedm.12228

Palermo, C., & Thomson, M. M. (2019). Large-scale assessment as professional development: Teachers' motivations, ability beliefs, and values. Teacher Development, 23(2), 192–212. https://doi.org/10.1080/13664530.2018.1536612

2018

Chen, J. (2018). KR-20. In B. Frey (Ed.), Encyclopedia of educational research, measurement, and evaluation. Sage Publishing.

Chen, J. (2018). Interstate School Leaders Licensure Consortium (ISLLC) standards. In B. Frey (Ed.), Encyclopedia of educational research, measurement, and evaluation. Sage Publishing.

Chen, J. & Perie, M. (2018). Comparability with computer-based assessment: Does screen size matter? Computers in the Schools, 35(4), 268–283. https://doi.org/10.1080/07380569.2018.1531599

Cui, Z., Liu, C., He, Y., & Chen, H. (2018). Evaluation of a new method for providing full review opportunities in computerized adaptive testing—Computerized adaptive testing with salt. Journal of Educational Measurement, 55(4), 582–594. https://doi.org/10.1111/jedm.12193

Conferences

Presentations

Conference paper and poster presentations by our staff.

2026

Huang, Y., Palermo, C., He, Y., Chen, T. (2026, April). Evaluating validity-based anchor responses as an external criterion for rater accuracy. Paper presented at the annual meeting of the National Council on Measurement in Education, Los Angeles, CA.

Huang, Y., Wilson, J., & Palermo, C. (2026, April). Using GPT-4 for automated essay scoring: Accuracy and fairness across student groups. Paper presented at the annual meeting of the National Council on Measurement in Education, Los Angeles, CA.

Palermo, C. (2026, April). Evaluating the generalizability of automated scoring models to a state-level assessment context. Paper presented at the annual meeting of the National Council on Measurement in Education, Los Angeles, CA.

2025

Beck, M. F., Liu, R., Chen, T., & He, Y. (2025, October). Efficacy of ad-hoc confidence measures and conformal prediction intervals in hybrid scoring. Paper presented at the Artificial Intelligence in Measurement and Education Conference, Pittsburgh, PA.

Chen, D., Huang, Y., Ji, X. & Guo, Y. (2025, April). Enhancing systematic reviews with large language models: Using GPT-4 and Kimi. eBoard presented at the annual meeting of the National Council on Measurement in Education, Denver, CO.

Chen, J., & Jiang, N. (2025, April). Exploring person-fit patterns in high-stakes alternate educational testing. Paper presented at the annual meeting of the American Educational Research Association, Denver, CO.

Chen, J., Jiang, N. & Huang, Y. (2025, October). Comparison of AI and human scoring on a visual arts assessment. Paper presented at the 2025 Artificial Intelligence in Measurement and Education Conference (AIME-Con 2025), Pittsburgh, PA.

Gui, Y. (2025, April). A Python-based automated test assembly (ATA) system. Paper presented at the annual meeting of the National Council on Measurement in Education (NCME), Denver, CO.

Gui, Y. (2025, October). Efficient AES: Dimensionality compression and distillation in transformer-based models. Paper presented at the Artificial Intelligence in Measurement and Education Conference (AIME-Con 2025), Pittsburgh, PA.

Huang, Y., Palermo, C., & Wilson, J. (2025, April). Exploring core intervention components in a randomized controlled trial on an automated writing evaluation system. Paper presented at the 2025 American Educational Research Association (AERA) Annual Meeting, Denver, CO.

Huang, Y., Wilson, J., & Palermo, C. (2025, April). Leveraging large language models and prompt engineering for automated essay scoring. Paper presented at the annual meeting of the National Council on Measurement in Education, Denver, CO.

Huang, Y. & Wilson, J. (2025, October). Evaluating LLM-based automated essay scoring: Accuracy, fairness, and validity. Paper presented at the 2025 Artificial Intelligence in Measurement and Education Conference (AIME-Con 2025), Pittsburgh, PA.

Palermo, C. (2025, April). Fostering self-regulated learning in writing: A standards-based analysis of MI Write. Paper presented at the annual meeting of the National Council on Measurement in Education, Denver, CO.

Palermo, C., He, Y., Justice, D., & Katula, P. (2025, October). Improving automated scoring accuracy through synthetic response generation and validation. Paper presented at the Artificial Intelligence in Measurement & Education Conference (AIME-Con), Pittsburgh, PA.

Wilson, J., Palermo, C., Deane, P., Ormerod, C., & Shermis, M. D. (2025, April). A framework to guide the development and evaluation of AI-based formative assessment systems. Paper presented at the annual meeting of the National Council on Measurement in Education, Denver, CO.

Wilson, J., Palermo, C., & Wibowo, A. (2025, October). Using AWE to measure writing growth among middle school ELs and Non-ELs. Paper presented at the Artificial Intelligence in Measurement & Education Conference (AIME-Con), Pittsburgh, PA.

2024

Huang, Y. & Wilson, J. (2024, April). The effects of an automated writing evaluation system on state test performance. Paper presented at the annual meeting of the National Council on Measurement in Education, Philadelphia, PA.

2023

Cruz Cordero, T., Wilson, J., Palermo, C., Eacker, H., Myers, M., Potter, A., & Coles, J. (2023, August). Writing motivation and ability profiles and transition after a technology-based writing intervention. Paper presented at the biennial European Association for Research on Learning and Instruction (EARLI) conference, Thessaloniki, Greece.

Cruz Cordero, T., Wilson, J., Palermo, C., Eacker, H., Myers, M., Potter, A., & Coles, J. (2023, April). Middle-school writing motivation: Profiles and transition in response to a technology-based writing intervention. Poster presented at the annual conference of the American Educational Research Association, Chicago, IL.

He, Y., & Chen, T. (2023, April). Using normalized theta score differences to evaluate equating with item parameter drifts. Paper presented at the annual meeting of the National Council on Measurement in Education, Chicago, IL.

Wibowo, A., Palermo, C., Vaughn, D., Justice, D. & He, Y. (2023, April). Combining linguistic features with deep neural network models to fine-tune response predictions for NAEP reading items. Paper presented at the annual meeting of the National Council on Measurement in Education, Chicago, IL.

Wilson, J., Palermo, C., Myers, M., Cruz Cordero, T., Eacker, H., Coles, J., & Potter, A. (2023, April). Impact of MI Write automated writing evaluation on middle grade writing outcomes. Paper presented at the annual meeting of the National Council on Measurement in Education, Chicago, IL.

Zhang, F., Wilson, J., Cruz Cordero, T., Palermo, C., Eacker, H., Myers, M., Coles, J., & Potter, A. (2023, April). Identifying predictors of middle school students' perceptions of automated writing evaluation. Paper presented at the annual meeting of the National Council on Measurement in Education, Chicago, IL.

2022

He, Y., Jing, S., & Lu, Y. (2022, April). A multilevel multinomial logit approach to bias detection. Paper presented at the annual meeting of the National Council on Measurement in Education, San Diego, CA.

Huang, Y., Potter, A., & Wilson, J. (2022, April). Teachers' perceptions of the validity of an automated writing evaluation system. Paper presented at the 2022 National Council for Measurement in Education (NCME) Annual Meeting, San Diego, CA.

Palermo, C. (2022, April). Examining hybrid automated scoring/handscoring results in a multi-state design. Paper presented at the annual meeting of the National Council on Measurement in Education, San Diego, CA.

2021

Cui, Z., Liu, C., & He, Y. (2021, June). Using machine learning to administer salt items in computerized adaptive testing. Paper presented at the annual meeting of the National Council on Measurement in Education (virtual).

Huang, Y., & Wilson, J. (2021, October). Exploring validity of automated essay scoring among English language learners. Poster presented at the 2021 National Council for Measurement in Education (NCME) Classroom Assessment Conference (virtual).

Huang, K., & Wilson, J. (2021, June). Using automated feedback to develop writing proficiency. Paper presented at the 2021 National Council for Measurement in Education (NCME) Annual Meeting (virtual).

Thacker, A., Word, A., Sinclair, A., Nash, B., & Chen, J. (2021, June). Moving bookmark standards setting from in-person to virtual: Best practices/lessons learned. Paper presented at the annual meeting of the National Council on Measurement in Education, online.

2020

He, Y., Wu, Y. F., & Tao, W. (2020, September). Comparing CTT postequating and IRT preequating in the embedded field-test model. Paper presented at the annual meeting of the National Council on Measurement in Education.

Murray, A., Daoust, C., & Chen, J. (2020, April). Validating tools for measuring Montessori implementation. Paper presented at the annual meeting of the American Educational Research Association, San Francisco, CA.

Wilson, J., Huang, Y., Beard, G., & MacArthur, C. A. (2020, April). A research-practice partnership examining the use of automated writing evaluation software: Effects on writing outcomes. Poster scheduled to be presented at the 2020 American Educational Research Association (AERA) Annual Meeting (conference canceled).

Wu, Y. F., He, Y., & Tao, W. (2020, September). Evaluating impacts on operational item performance in the embedded field-test model. Paper presented at the annual meeting of the National Council on Measurement in Education.

2019

Cui, Z., Liu, C., & He, Y. (2019, April). On administering salt items in computerized adaptive testing with salt. Paper presented at the annual meeting of the National Council on Measurement in Education, Toronto, Canada.

Daoust, C., Murray, A., & Chen, J. (2019, March). A reexamination of implementation practices in Montessori early childhood education. Paper presented at The Montessori Event, Washington, DC.

Murray, A., Chen, J., Daoust, C., & Amos, A. (2019, April). Dimensions of fidelity in a constructivist classroom. Paper presented at the annual meeting of the American Educational Research Association, Toronto, Canada.

Wilson, J., Beard, G., Huang, Y., & MacArthur, C. A. (2019, September). A research-practice partnership aimed at improving writing outcomes via implementation of automated writing evaluation. Poster presented at the 2019 National Council for Measurement in Education (NCME) Classroom Assessment Conference, Boulder, CO.

Wilson, J., Huang, Y., Beard, G., & MacArthur, C. A. (2019, December). Supporting writing instruction and writing outcomes in the elementary grades using automated writing evaluation software: Results from a district-wide implementation. The 69th Literacy Research Association Annual Conference, Tampa, FL.

2018

Chen, T., Tao, W., & Gao, X. (2018, July). Evaluating item position effects on scrambled form pre-equating. Paper presented at the annual meeting of the International Test Commission Conference, Montreal, Canada.

Wang, W., Zheng, Z., & Chen, J. (2018, October). Clustering students in a state classroom assessment system: Exploring the usages for classroom assessment. Paper presented at the National Council on Measurement in Education Special Conference on Classroom Assessment, Lawrence, KS.

2017

Chen, T., Huang, C. H., & Liu, C. (2017). An imputation approach to handling incomplete computerized tests. Paper presented at the annual meeting of the International Association of Computerized Adaptive Testing, Niigata, Japan.

Fang, Y., Lu, Y., & He, Y. (2017, April). Can subtest equating borrow information from the full test? Paper presented at the annual meeting of the National Council on Measurement in Education, San Antonio, TX.

He, Y., & Yi, Q. (2017, April). Impact of item parameter drift on mixed-format tests. Paper presented at the annual meeting of the National Council on Measurement in Education, San Antonio, TX.

Collaborate with us

Interested in our research?

Our psychometricians and scientists partner with clients and researchers to advance the science of measurement. Reach out to learn more.