Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Speaking from the Devil's Advocate side, as I work in Higher Education analytics:

First: this data is, by the vast majority of institutions, used very respectfully in terms of privacy rights and it is ABSOLUTELY never sold. Third part vendors that provide hosted systems are REQUIRED to affirm that they pass nothing on to any other party whatsoever. (I should note that this is the case in public and traditional private non-profit schools. For-profit schools are much less scrupulous and should not be regarded as trustworthy in this respect)

Beyond this, students are apprised of these uses of data with the ability to opt out. This is done as upfront as we can, very "in your face", though the tendency to click "agree" to something without reading, even very short and to the point, is of course always an issue.

Finally, the data collected is extremely useful. I cannot emphasize that enough. It is extremely useful in helping students be more successful. And it's not used in terms of surveillance. It's used pretty much exclusive in understanding the patterns the coincide with success and those that coincide with failure, dropout, academic struggling, and so on. In short, it is used to proactively intervene with students at risk rather than wait for it to be too late, e.g., final grades are in and no tutoring intervention is possible. A recent use had me using Xgboost to examine patterns of LMS usage to determine students at risk of early drop out. It worked, and we can now identify about one third of early drop outs with practically no effort and almost no false positives and work with those students to address there problems before they go too far.

I'm sure privacy maximalists will not be satisfied with the above, but really, within the realm of data comprehensive I idividual data collection, I'm not aware of any business, much less an entire industry, that treats such data as responsibly as this, and that really is concerned with using it only to further the best interests of it's constituents. Obviously I'm biased. That probably means things are not as rosy as I envision. But I think it is nonetheless an important and valid synopsis of the issue.



Just because the institution does not use the data irresponsibly doesn't mean the service provider will not. Not to mention third parties which could potentially gain access to this data directly or indirectly by listening for the check in signal devices send.

Per the article, this is not being used to help the students succeed, instead it is being used to damage their grade for a lack of attendance. I think most reasonable people would agree attendance is not a reliable metric by which to base a students grade. The exception being discussion based or interactive courses like debate or public speaking.


We have extremely stringent security checklists for 3rd part vendors, as I alluded to in my post. Both the vendors, and any of the vendor's vendors, must be able to demonstrate that data is not used for any purpose but the school's. As for attendance, automated attendance tracking has been around for decades in the form of swipe cards, and simply automated a process professors undertook manually. No data is captured that would not have been in the analog world, and professors are still free to have no attendance policies at all schools I have spoken to about it. Granted, that may only be a sample size of 100 or so.


I work in higher ed and we use a student success analytics platform. Where I see this going wrong is administrators want a system to do everything. We want to identify the at-risk student, nudge them automatically, and have a successful outcome. Schools don't want to invest in people to actually make outreach or connection with students because that is hard. As a result, we generate a ton of data with little to show other than numbers.

I'd be very curious to learn about your LMS usage project. Are you saying you identified at-risk students quickly enough to intervene at a point in the term where it made a difference between passing and failing and that the school successfully intervened in those cases?


We're not perfect, but as we have brought on analytics systems and engaged in institutions drive projects like my use of Xgboos with LMS data, we have invested heavily in additional advising resources because we recognize that data is irrelevant if you don't have people that will take action on it. I can make data driven predictions all day long about students at risk. What matters is what do we do about it? I've spoken with schools that thought buying an analytics platform was all they had to do. They have universally been disappointed with their lack of results.


That's great, thanks.

I get frustrated as a technical staff person, who personally enjoys being on the cutting edge, often being the one in meetings saying 'data is great but how will we actually use it?'.

Honestly, I'll probably retire form higher ed in a few years and go work for an outfit like yours.


Hah, actually the "outfit like mine" is a university that has relatively recently come to some very good conclusions on this sort of thing. 5 or 6 years ago... Not so much. But one nice thing about working for an institution like this is that mostly people are looking to do the right thing for students, and don't just look at the thing as some commercial revenue extraction machine.


Ahhhh now I get it. It's so rare for institutions to do internal projects like this I didn't even consider it to be anything but a commercial revenue extraction machine.


> Beyond this, students are apprised of these uses of data with the ability to opt out.

What happens if they opt out?

And can they opt out of assessment by default, while still allowing data collection? That is, could the data be collected, but go into a secure archive, which only the student could access? So if the student isn't doing well, they could release selected data to advisers.


Unfortunately it's not as granular as that, so there is still lots of room for improvement. There are some specific things that can be opted out of as required by FERPA, but otherwise it's a bit all or nothing.


I think I understand the rationale behind this fairly well, and find it sensible to a certain degree.

With that said, I find a particular part of your explanation more intellectually dishonest than the anal probe comment. That's not to say you're being flippant as perhaps that commenter was.

To claim that you're not using this data in terms of surveillance is utterly nonsensical. If you can demonstrate to me how you're not actively conducting surveillance, I'm open to hearing it.

Privacy trade offs may provide measurable gains, but I've yet to come across even attempts at quantifying the loss of individual privacy. You can show asymptotically improving trends, but the costs seem to remain blank, unknown, and I think that is what our flippant friend was trying to get at.


Thanks for the respectful and nuanced take.

I am somewhat concerned with where this stuff leads, not because I assume it's a one step process between education analytics and oppressive dystopia, but precisely because the use cases you describe seem like great ideas.

I agree that minimizing drop-out rates is important, and that we should avail ourselves of whatever tools we can to do it. What you're describing does not bother me, and seems like a reasonable use of the data.

So why not take it a step further? Let's authorize a limited release of student-identifying data to government social work agencies (like the DHS in the United States), so that concerning student performance data could be correlated with whether a student's family was on income/food/housing assistance. Heck, to maximize the signal, we could even correlate it with families who aren't on assistance but whose homes are likely to be foreclosed upon (leading to a likely student move, possibly changing schools--a very disruptive educational experience whose impact can be mitigated with careful preparation). Why not correlate with criminal records, to help promote after-school programs to students who would otherwise be returning to troubled or dangerous homes?

To be clear, I'm bothered because those things all seem like good ideas. I think they'd probably work, if done properly. And would probably help a lot of people. The systems we have currently to try to accomplish those things (overworked social workers, underfunded public-education and after-school programs) aren't succeeding.

...they'd also result in a de facto surveillance state.

If that data were compromised by or sold to an outside (nongovernment) entity, the damage due to the breach would be very severe, from inability to get financing (because lenders know details about your life that may not be reflected on your credit score) to individual persecution (e.g. employer discrimination against family members whose leaked data indicated their financial/housing situation may put their long-term employment at risk).

If the government programs using that data were "compromised" (i.e. staffed, perhaps via an electoral sea-change, with discriminatory people), the damage would be even worse. School funding allocation could suffer because someone at the top decided "Our data shows that so many of the students at $school_x have so much else going against them in their lives--why bother? Starve that school for funding and let it wither; that money would be better allocated to schools attended by our kids/kids of different races/kids with higher income." Additionally, correlations could be dishonestly interpreted by bigoted officials ("the algorithm says that female/male/dyslexic/$race_x kids are a lost cause!").

There are other, logistical issues. If school staff (i.e. counselors) have direct access to all this information, could they discriminate against kids and their families if bigoted? If school staff didn't have direct access, will the right conclusions be drawn by distant analysts (or worse, algorithms) without the personal context of counselors and teachers? A failure to consider all those factors could result in IEP-like things that harm, rather than help (consider: "You talked back to counselor Bob, so he won't authorize you for additional meal points or recommend you for after-school tutoring in an area where you're struggling", or "the DOE algorithm decided you need to take more English classes even though counselor John has been working on getting you individual reading time because he knows there's a bullying risk from other kids in the English program").

Many of these things are already happening, in varying degrees. I worry that making ubiquitous data available could make them worse, if it is not done carefully.

I am also sensitive to the benefits that education analytics (and more pervasive, cross-cutting data availability for government use across the board) can provide. I think the answer is to proceed carefully (i.e. the opt-out buttons you discuss, and careful data anonymization whenever it makes sense to do so) and to insist on comprehensive regulation regarding data exchange between government programs, and between those programs and non-government entities.

EDITS: explained logistical issues. Typos.


You're absolutely right, data could be used for these other things that, while potentially beneficial to society and individuals, carry further potential for misuse as well.

To address misuse in general: That is always a problem, with any technology, tool, data... anything. It's important that we don't discard useful tools due to potential misuse though. Instead, it's of the utmost importance that, at the outset, issues of misuse be enumerated, discussed, mitigation strategies thought out, how to accurately talk about data gathered, the insights obtained, what they can and cannot be used for from a statistical analysis (correlation != causation etc. type of thing) and so on. Let me talk a little about what we did at my institution along these lines. It wasn't perfect, we could have done better, but in the outlines I think it formed a pretty good starting place for a template of how society should engage in conversations about issues that arise from the current revolution in our ability to gather data on individuals:

In 2014 we started using an LMS system that made data collection on easy. I was given sole access to that data to begin exploring and cataloging the available data being collected, it's possible uses, potential points of concern. I had sole access, but those things I was looking for were done in intense conversation with administrative and faculty colleagues, including the university executive board and the faculty union and faculty senate along with conversations with students. During this process, no else had the level of unfettered access that I did (I am a data analyst) not even the system administrators. (Technically they could have granted themselves access though, but if they didn't go through the proper channels, such a thing is a firing offense.)

This went on for about three years, all the while I was able to accumulate data and evaluate it against actual student performance and validate its utility and continue to report back on my findings to further the discussions going on. When that stage of actually using the data to drive conversations ended, I took a step back and simply spoke at gatherings every once in a while to go over what I'd done, findings, etc. After that, for the next two years, it was in the hands of the administration working up policies in conversation with faculty surrounding expectations for responsible use of this information. To say there were disagreements would be an understatement. But over time tentative points of compromise were reached, and for the past 2 years were have begun carefully rolling out initiatives that utilize all of this. We are cautious. We are open. We continue to have conversations with valid concerns on all sides that we seek good compromises for.

Again, the above wasn't perfect, but it's a good outline of how responsible use of data should be approached.


Is it open source / free software?

These things should not be allowed to be proprietary software.. if I attend a public institution that surveils me, I want the surveilance software to be fully open source and I want to be able to know all the data points you would have on me.

Then, I would like to decide for myself whether I want to participate or not.


“ respectfully in terms of privacy rights and it is ABSOLUTELY never sold.”

Could it be hacked? How good are university IT security operations?


At my school, they're surprising good. I mean, the systems aren't in an air gap, but we use 2FA, requests for data access are scrutinized, and not from a bureaucratic perspective but from a conversational perspective: "Let's talk about what you're trying to do with this data, what's the minimum needed for your goal? Is on-going access really necessary or can we provide a single extract that you access via a secure environment?" That sort of thing. Infrastructure security isn't something I'm involved with, but we do have a robust security team that works on it, that I engage with when working through vendor checklists and other security matters, and they are up to date on the field and conscientious to the point of (reasonable) paranoia.

We're also not a high-profile target: Most hackers looking to crack a system holding personal data seem most concerned with high value targets with millions, not thousands of records. That doesn't stop our security team from treating the possibility with the utmost seriousness though.

However, the above applies to where I work. This is not universally the case. I think most schools are increasingly aware of the potential, are take minimally necessary steps to make sure they aren't wide open, but there are simply too many small schools with 1,000 to 5,000 students that simply don't have the scale needed to really devote resources to this. Again, the low-value target helps here, but it's not great.


> We're also not a high-profile target

Really now? Have you checked your school on Shodan recently?

Long story short, there's enough people doing hacking, with varied enough interests. You don't know what makes someone else tick.


[flagged]


That is an intellectually dishonest extension of this to the absurd. If you wish to discuss the merits or faults of this I will do so. This is not some clear cut issue. I know there are privacy concerns. We try to handle them as transparently and honestly as possible. In that effort, we save students millions of tuition dollars and years of their lives. I am receptive to arguments that the privacy tradeoff isn't worth it, we have that discussion every time we examine possible data sources and have decided not to use some. But your comment is beneath any ability to have a reasonable discussion on the topic.


That's a super uncharitable interpretation of the GP's argument.

Privacy is a spectrum. By your "logic", should students' entire identities be concealed from education personnel? Of course not. That's just as ridiculous as what you proposed.


Arguably, yes. As covered in Stephenson's Dodge. Students are fully anonymized, in meatspace as well as in data. So there's no possibility that anything except performance affects assessment.


Stephenson managed to present both a Utopian and Dystopian view of the future at the same time... I'm not sure his work there is very instructive in this matter. But the thing is, factors other than quantifiable performance do impact assessment. As just one small example, from what I have seen in analyzing the data, the best performing students at some of the worst inner city schools tend not to be nearly as prepared for college as comparable students from better school districts. In short, there quantifiable performance, anonymous from further detail and context, would not accurately predict their academic performance. At the same time, these students typically have overcome many more obstacles to get to where they are, and while they do not always get the same grades, their determination often sees them through to the end where their counterparts may not. Context matters, and the more context you get, the less anonymous things become. That's the problem with either extremist view of total open "privacy is dead" views or total privacy maximism. They seek to avoid the hard work of dealing with the ambiguity in favor of black/white views. Because it is hard work. It involves messy questions. Do we save students hundreds of thousands of dollars and years of their lives by tracking details of course interactions and systems usages? A noble goal, put that way it's hard to say no. But what if the student doesn't want it? Do we give up on it completely? What if we can balance it somehow with informed consent, is that sufficient? If not, is there something we can do that doesn't sacrifice things that are beneficial for students but also allows each person to fully control their data footprint? If we go ahead with any of this, what further monitoring might we then justify? Is it a slippery slope or do we have checks & balances in place to make sure data capture & usage doesn't leap ahead of carefully considered policy & consequences? None of the above are easy. You can ignore them all if you choose a black or white view of things. But either way you give up quite a bit, and whether you choose black or whether you choose white, society will still exist in unending shades of gray.


That's true about the novel [Fall, not Dodge]. But not, I think, about the idea of students being anonymous.

I mean, it's not that Sophia remained entirely anonymous. But she did retain full control over who had access to what. She was anonymous in applying for an internship. But as a grad student, she got lots of expert coaching.


On the contrary, it's a concise rejection of the worldview that privacy means privacy from others, not privacy from me, and that students are interchangeable units subject not to any consideration of ethics but only performance.


Where is that worldview evidenced in the post you were responding to?

The poster seemed extremely cognizant of privacy concerns, and was defending a limited collection of specific data under very narrow circumstances for similarly delimited purposes. Put another way, saying "I think some specific data collection is ethical" is not the same as saying "all data collection I do is ethical, others may not be ethical".

They also seemed to implicitly reject the students-are-interchangeable idea: the whole point of gathering this data is to determine if students (as individuals, classes, or entire schools) are struggling or in need of help. That's the opposite of interchangeable; that's finding suffering outliers.

I'm harping on it 'cause I think I may be closer to your position on this issue than I am to the original poster's, but ... you're making extremist straw men, talking about feeding trays, and all but reciting the lyrics from the Pink Floyd song. I think that's making pro-privacy positions look bad.

EDIT also what about my first point: is any data collection on students ethical? Names? Addresses? Test results? Or should teachers only operate behind double-blind, voice-and-handwriting-obfuscated screens from their students? See? I can make straw men, too.


The OP's main concern about privacy was that "third parties" might get the data. There is no recognition of the fact that a student might deserve or expect privacy from the school, only from some nebulous Other who is certain not to treat it with the great respect sure to be afforded it by the Benevolent and Definitely Not-For-Profit university. There is no question of whether this totalizing monitoring of a life is an affront to dignity or whether it normalizes authoritarian power structures, only that it is "extremely useful", that is, it subjects pure quantities to optimization.

There is already a lower bound on the level of acceptable data collection set by normal human interaction. There does not however appear to be any upper bound whatsoever suggested by the argument in the OP. As long as the data is not sold and it continues to be useful (ie. the Mathusian race has not yet reached the bottom), no aspect of human life appears to be too sacred to be put into into the optimizer.

(I am not wott, in case you did not notice.)


That fact that it's useful doesn't mean there isn't a human face to it. Look at the words I use. I don't say "optimizing outcomes" or some other dehumanizing turn of phrase. When I said it was "extremely useful", I said that in the context of "helping students be more successful" To interpret my words here as "quantities to optimization" is to impose your own views on these practices as a whole instead of to the specific example I spoke of, which does not do so. If you look at one of my other comments, we spent years engaging with students, administrators, and faculty before we put into place initiatives that utilize this sort of data.

As for recognizing the fact that students might expect or deserve privacy, and the affront to dignity, etc., I was quite clear that we are as transparent an upfront as we can be about this, and students have the ability to tell us "no". Beyond that, the concept of "affront to dignity" on the matter is a personal opinion on the topic, not an axiomatic ethical marker. Similarly "expectation of privacy" is not an unmovable value. My school, in part because of a years-long careful deliberation on the topic, was behind the times in this area. As a result, I cannot tell you how many times students speak to advisors and are actually frustrated that the advisor does not know some of these things about the student, that the advisor has to ask if the student has been attending classes, what their midterm grade is, etc. The students expect us to have all of this data at our fingertips because they expect a seamless, friction-less level of service.

Maybe you'd say "that's just because they're used to so many organizations collecting data," and sure, you'd be right (my apologies if the the words I put in your mouth aren't one's you'd say or agree with) But here's the thing: If it's what they expect, what they want, and done responsibly, what is the problem? To contradict the students' attitudes puts you in a position of telling other people that what they want is wrong, and that they should want other things that you think are best for them. Again, my apologies if that doesn't correspond with your own ideas. It is, however, not an uncommon stance, so I think the point still stands.


I don't expect a university advisor to have surveillance data on me, I expect them to be busy with their own work, and that I can come to them for assistance when I need. It's up to the student to communicate the situation.. this should be a human interaction and not automated.


It's a concise reduction to the absurd that doesn't leave room for conversation. It's a flat statement of "your wrong" that adds nothing to the conversation, with the subtext of "I'm too lazy to actually put together an argument with a claim and a few premises and a bit of explanation to flesh things out".

It's the opposite of constructive. It's combative and inflammatory. HN, unlike many other venues for discussion, is a halfway decent community precisely because such comments are in the minority and tend to be quickly downvoted.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: