I love JupyterHub, but we've hit some real headaches in having it scale. At Virginia Tech, we have an introductory course for non-computing majors where students are using Jupyter through JH. At around the 70-student mark, we have performance issues. Considering that the course is eventually meant to scale much farther (hundreds of students), we're not really sure how we can make further progress with our current resources. I hope this new version has some performance enhancements (though I don't see any in the changelog). Last I talked with anyone about this was in the NBGrader project[1], where other schools were hitting walls with scaling.
Give us some details about your setup. Are you running JH on the cloud somewhere? Or is this running on a single box at VT? We’ve got JH running on a Kube cluster with persistent storage on the NSF Jetstream cloud [1] (which you could qualify for free resources being from VT) . This set up should theoretically address the scalability concerns you are having. See the work Andrea Zonca and I have been doing [2, 3, 4, 5]. Contact me for additional details.
I'm not as up on the details as you'd probably want, but my understanding is that we're running it on a virtual server with a nice chunk of RAM and CPU locally at VT. I don't think we want to be at all reliant on external servers - this is FERPA protected data. Plus, the long term goal was to find a solution that other schools could adopt without being an R1.
Are you running into scaling issues because all students are on a single box? I think the JupyterHub proxying service should be able to handle way more than 70 users, but I could see 70 users on a single machine being problematic. I'm doing a lot of work around JupyterHub - happy to help out if I can, I can be reached at hugo@saturncloud.io
Well, it's a virtual server, so I'm not sure how that plays with things. I supposed getting more servers in play would help, but that feels more like throwing hardware at a software problem. It doesn't feel like JH should have this much overhead - is the kernel really doing so much work?
Depends on what your students are doing right? Also I'd bet your running out of ram before CPU. You should try doing the same workload without JH involved - I'd bet you'd still run out of ram. Also - I'd throw hardware at software problems all day long if I could. Hardware is cheap compared to developer time
[1] https://github.com/jupyter/nbgrader/issues/530#issuecomment-...