Generative AI as a Tool for Creating Qualitative Rubrics
Type
conference contribution
Date Issued
2024-04-12
Author(s)
Abstract
The proposed proposal addresses three problem areas and presents practice-oriented solutions for university teaching involving generative AI. These areas are:
1. the development of grading in times of AI
2. assistance through systematized grading schemes (rubrics)
3. the assistance of various AI systems in the creation of rubrics
The input will address these three problem areas and conclude with a case study that was successfully implemented at the University of St.Gallen (HSG).
Certain types of exams have come under pressure due to the emergence of generative AI sys-tems. The ability to generate texts seemingly at the "push of a button" means that the classic seminar paper or term paper in particular is increasingly under pressure to legitimize itself. But the same applies to presentations, essays and reflection reports (Fyfe, 2023). In fact, from the perspective of didactic teaching development, more or less all forms of examination must be put to the test in view of the effects of generative AI applications. Universities still do not seem to be fully aware of this problem: The current, and only beginning, transformation will have an impact on the examination process in the university context in general (Getto et al., 2018a, 2018b). Certain legal adjustments are necessary for this, but not sufficient to do full jus-tice to the phenomenon.
The University of St. Gallen has developed an initial classification grid for this purpose, which can also be presented and discussed during the workshop.
Rubrics are one way of making grading in the digital age (Hauck-Thum & Noller, 2021) trans-parent and fair. At least if they are made known to students as assessment schemes. They can therefore help not only to eliminate any uncertainties that generally exist in written work, but also to address the challenges indicated (Mahmood & Jacobo, 2019).
In general, rubrics are an attempt to create a rational and comprehensible evaluation grid that is also based on the conventions of the subject area. However, the parameters chosen are not arbitrary, but share certain assessment categories such as depth of content presented, structure of the work, style and scientific language use and formality. Such a scheme has certain ad-vantages, but also disadvantages (Jahn & Cursio, 2021). One of the advantages is that students receive guidance on the assessment, which is often seen as arbitrary. Furthermore, academic (and non-academic) writing can be practiced using predetermined criteria to provide students with the skills that are essential for a potential future academic career.
The undeniable disadvantages include, on the one hand, the possible lack of clarity and, on the other, the effort involved in creating a rubric. In the former case, it can be stated that scale levels in rubrics often only provide generally valid parameters ("the student has fully/less/not at all grasped the content..."). Such assessments nevertheless lead back to a subjective evalua-tion. Experience, particularly with regard to the assessment of students' skills and knowledge, therefore remains indispensable on the part of lecturers.
But any creative work, and written assignments at a university are undoubtedly creative work, harbors certain uncertainties in an assessment situation that can be explained by certain aca-demic and non-academic influences (Kant, 2009). AI applications cannot help with this fun-damental problem.
However, they can help in the categories relating to effort. In our experience, the use of AI tools can provide added value here. In order to be able to respond to the problem of fuzziness, rubrics must be geared towards the learning objectives and learning taxonomies that are at the beginning of every examination design.
Only through this connection can the required rational depth in the assessment be achieved; i.e. only here is it ensured that the test has also tested what it is supposed to test (Biggs, 1996). This means that rubrics, if they are to have a didactic meaning, must be created individually for each course or for each question.
The joint proposal should therefore show the possibilities of AI-supported creation of rubrics, name the advantages and disadvantages of the method and provide a best-practice example for use. During our slot, we will also be happy to engage in practical work with attendees who want to use various AI models for practical application in their course or program.
In practical terms, several departments of the HSG's Vice President's Office for Teaching and Learning, namely the Teaching Innovation Lab, the Writing Lab and the Center for Curriculum and Teaching Development, have investigated which freely available AI applications are cur-rently particularly suitable for creating rubrics. This took place in a 90-minute workshop with various lecturers and different levels of experience in dealing with AI and rubrics. The chat-bots from OpenAI versions 3.5 and 4, Google Bard and Microsoft's Co-Pilot were tested.
The results were thus collected qualitatively and can be verified again before the conference. Until this verification, no concrete results can be given here, but there was a clear trend that showed ChatGPT 3.5 to be slightly ahead of the other LLMs mentioned.
The specific questions that were asked and the extent to which these results are comparable will be presented at the conference.
The contribution will therefore be a practical demonstration of the possibilities and limitations of AI tools in the creation of rubrics. Furthermore, the research is intended to stimulate the exploration of the possible applications and limits of AI currently being discussed in teaching in order to provide possible good-practice examples in dealing with such developments.
Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32(3), 347-364. https://doi.org/10.1007/bf00138871
Fyfe, P. (2023). How to cheat on your final paper: Assigning AI for student writing. AI & SOCIETY, 38(4), 1395-1405. https://doi.org/10.1007/s00146-022-01397-z
Getto, B., Hintze, P., & Kerres, M. (2018a). Digitalisierung und Hochschulentwicklung. Waxmann Verlag.
Getto, B., Hintze, P., & Kerres, M. (2018b). (Wie) Kann Digitalisierung zur Hochschulentwicklung beitragen?
Hauck-Thum, U., & Noller, J. (2021). Was ist Digitalität? https://doi.org/10.1007/978-3-662-62989-5
Jahn, D., & Cursio, M. (2021). Kritisches Denken. https://doi.org/10.1007/978-3-658-34985-1
Kant, I. (2009). Kritik der Urteilskraft (H. F. Klemme & P. Giordanetti, Eds.). Felix Meiner Verlag. https://doi.org/10.28937/978-3-7873-2069-1
Mahmood, D., & Jacobo, H. (2019). Grading for growth: Using sliding scale rubrics to motivate struggling learners. Interdisciplinary Journal of Problem-Based Learning, 13(2).
1. the development of grading in times of AI
2. assistance through systematized grading schemes (rubrics)
3. the assistance of various AI systems in the creation of rubrics
The input will address these three problem areas and conclude with a case study that was successfully implemented at the University of St.Gallen (HSG).
Certain types of exams have come under pressure due to the emergence of generative AI sys-tems. The ability to generate texts seemingly at the "push of a button" means that the classic seminar paper or term paper in particular is increasingly under pressure to legitimize itself. But the same applies to presentations, essays and reflection reports (Fyfe, 2023). In fact, from the perspective of didactic teaching development, more or less all forms of examination must be put to the test in view of the effects of generative AI applications. Universities still do not seem to be fully aware of this problem: The current, and only beginning, transformation will have an impact on the examination process in the university context in general (Getto et al., 2018a, 2018b). Certain legal adjustments are necessary for this, but not sufficient to do full jus-tice to the phenomenon.
The University of St. Gallen has developed an initial classification grid for this purpose, which can also be presented and discussed during the workshop.
Rubrics are one way of making grading in the digital age (Hauck-Thum & Noller, 2021) trans-parent and fair. At least if they are made known to students as assessment schemes. They can therefore help not only to eliminate any uncertainties that generally exist in written work, but also to address the challenges indicated (Mahmood & Jacobo, 2019).
In general, rubrics are an attempt to create a rational and comprehensible evaluation grid that is also based on the conventions of the subject area. However, the parameters chosen are not arbitrary, but share certain assessment categories such as depth of content presented, structure of the work, style and scientific language use and formality. Such a scheme has certain ad-vantages, but also disadvantages (Jahn & Cursio, 2021). One of the advantages is that students receive guidance on the assessment, which is often seen as arbitrary. Furthermore, academic (and non-academic) writing can be practiced using predetermined criteria to provide students with the skills that are essential for a potential future academic career.
The undeniable disadvantages include, on the one hand, the possible lack of clarity and, on the other, the effort involved in creating a rubric. In the former case, it can be stated that scale levels in rubrics often only provide generally valid parameters ("the student has fully/less/not at all grasped the content..."). Such assessments nevertheless lead back to a subjective evalua-tion. Experience, particularly with regard to the assessment of students' skills and knowledge, therefore remains indispensable on the part of lecturers.
But any creative work, and written assignments at a university are undoubtedly creative work, harbors certain uncertainties in an assessment situation that can be explained by certain aca-demic and non-academic influences (Kant, 2009). AI applications cannot help with this fun-damental problem.
However, they can help in the categories relating to effort. In our experience, the use of AI tools can provide added value here. In order to be able to respond to the problem of fuzziness, rubrics must be geared towards the learning objectives and learning taxonomies that are at the beginning of every examination design.
Only through this connection can the required rational depth in the assessment be achieved; i.e. only here is it ensured that the test has also tested what it is supposed to test (Biggs, 1996). This means that rubrics, if they are to have a didactic meaning, must be created individually for each course or for each question.
The joint proposal should therefore show the possibilities of AI-supported creation of rubrics, name the advantages and disadvantages of the method and provide a best-practice example for use. During our slot, we will also be happy to engage in practical work with attendees who want to use various AI models for practical application in their course or program.
In practical terms, several departments of the HSG's Vice President's Office for Teaching and Learning, namely the Teaching Innovation Lab, the Writing Lab and the Center for Curriculum and Teaching Development, have investigated which freely available AI applications are cur-rently particularly suitable for creating rubrics. This took place in a 90-minute workshop with various lecturers and different levels of experience in dealing with AI and rubrics. The chat-bots from OpenAI versions 3.5 and 4, Google Bard and Microsoft's Co-Pilot were tested.
The results were thus collected qualitatively and can be verified again before the conference. Until this verification, no concrete results can be given here, but there was a clear trend that showed ChatGPT 3.5 to be slightly ahead of the other LLMs mentioned.
The specific questions that were asked and the extent to which these results are comparable will be presented at the conference.
The contribution will therefore be a practical demonstration of the possibilities and limitations of AI tools in the creation of rubrics. Furthermore, the research is intended to stimulate the exploration of the possible applications and limits of AI currently being discussed in teaching in order to provide possible good-practice examples in dealing with such developments.
Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32(3), 347-364. https://doi.org/10.1007/bf00138871
Fyfe, P. (2023). How to cheat on your final paper: Assigning AI for student writing. AI & SOCIETY, 38(4), 1395-1405. https://doi.org/10.1007/s00146-022-01397-z
Getto, B., Hintze, P., & Kerres, M. (2018a). Digitalisierung und Hochschulentwicklung. Waxmann Verlag.
Getto, B., Hintze, P., & Kerres, M. (2018b). (Wie) Kann Digitalisierung zur Hochschulentwicklung beitragen?
Hauck-Thum, U., & Noller, J. (2021). Was ist Digitalität? https://doi.org/10.1007/978-3-662-62989-5
Jahn, D., & Cursio, M. (2021). Kritisches Denken. https://doi.org/10.1007/978-3-658-34985-1
Kant, I. (2009). Kritik der Urteilskraft (H. F. Klemme & P. Giordanetti, Eds.). Felix Meiner Verlag. https://doi.org/10.28937/978-3-7873-2069-1
Mahmood, D., & Jacobo, H. (2019). Grading for growth: Using sliding scale rubrics to motivate struggling learners. Interdisciplinary Journal of Problem-Based Learning, 13(2).
Language
English
Event Title
Seamless Learning Conference 2024
Event Location
WU Vienna
Event Date
12.04.2024
Contact Email Address
sebastian.meisel@unisg.ch
File(s)![Thumbnail Image]()
Name
Generative AI as a Tool_Seamless Learning Conference.pptx
Size
1.8 MB
Format
Microsoft Powerpoint XML
Checksum (MD5)
5f4d406416d747dc03334602c5196a85