I am developing a smart RSS reader that ingests roughly 1000 articles a day and selects roughly 300 to show me. The current classification workhorse works on a miniLM embedding which is also used for clustering (unlike every other document clustering system i’ve seen, this one really works)
The performance of the classifier is limited by the fickleness of my judgements, I am thinking about making it into a bookmark manager, an image classifier, something that can sort through 5000 search results, and a workflow engine.
The performance of the classifier is limited by the fickleness of my judgements, I am thinking about making it into a bookmark manager, an image classifier, something that can sort through 5000 search results, and a workflow engine.