Article ID Journal Published Year Pages File Type
344277 Assessing Writing 2013 15 Pages PDF
Abstract

In this paper, we provide an overview of psychometric procedures and guidelines Educational Testing Service (ETS) uses to evaluate automated essay scoring for operational use. We briefly describe the e-rater system, the procedures and criteria used to evaluate e-rater, implications for a range of potential uses of e-rater, and directions for future research. The description of e-rater includes a summary of characteristics of writing covered by e-rater, variations in modeling techniques available, and the regression-based model building procedure. The evaluation procedures cover multiple criteria, including association with human scores, distributional differences, subgroup differences and association with external variables of interest. Expected levels of performance for each evaluation are provided. We conclude that the a priori establishment of performance expectations and the evaluation of performance of e-rater against these expectations help to ensure that automated scoring provides a positive contribution to the large-scale assessment of writing. We call for continuing transparency in the design of automated scoring systems and clear and consistent expectations of performance of automated scoring before using such systems operationally.

► The e-rater features and model types are described. ► The model building procedures and evaluation criteria for e-rater are described. ► The potential uses of e-rater for large-scale assessments of writing are discussed.

Related Topics
Social Sciences and Humanities Arts and Humanities Language and Linguistics
Authors
, ,