Anpther issue is one of copyright: obviously the student is the author. And we all know that the ML scoring subcontractor is keeping copies, with human ratings for later retraining purpouses.
At the time the student takes the test, he should be prompted with the informed choice asking him to grant either 1) no license to keep a copy for training purpouses 2) a non-exclusive license, and the website where he can get a copy of his own essay 3) a public domain license, again with the relevant domain linked so he can find his own and other's essays. 4) any of the above or other as a function of the resulting grade!
At the same time he should also specify his desire for or against attribution, again probably best as a function of the resulting grade. And under what moniker he wishes this contribution to exist.
These options to be filled out during exam time should have no default options (no opt-out), and preferably should be standardized by the community and lobbied for to be mandatorily enforced at state or federal level (forcing examinations to present the student with an informed choice)
A public dataset of legally obtained essays (without scores or names) would already be a very important first step to invite others to make actual performant ML grading systems.
I don't believe the current datasets in these "non-profit" organizations actually comply with copyright law, organizations who don't profit from the grading service towards the state, but do provide a stable ML job on the income from charging the people with financial means to test submissions, enabling a stealth class based society.
At the time the student takes the test, he should be prompted with the informed choice asking him to grant either 1) no license to keep a copy for training purpouses 2) a non-exclusive license, and the website where he can get a copy of his own essay 3) a public domain license, again with the relevant domain linked so he can find his own and other's essays. 4) any of the above or other as a function of the resulting grade!
At the same time he should also specify his desire for or against attribution, again probably best as a function of the resulting grade. And under what moniker he wishes this contribution to exist.
These options to be filled out during exam time should have no default options (no opt-out), and preferably should be standardized by the community and lobbied for to be mandatorily enforced at state or federal level (forcing examinations to present the student with an informed choice)
A public dataset of legally obtained essays (without scores or names) would already be a very important first step to invite others to make actual performant ML grading systems.
I don't believe the current datasets in these "non-profit" organizations actually comply with copyright law, organizations who don't profit from the grading service towards the state, but do provide a stable ML job on the income from charging the people with financial means to test submissions, enabling a stealth class based society.