worked on a little trivia question generator last night with astra. we built a good pipeline (weighted sample of wiki pages, feed a few dozen to small models, bigger model downselects), but no model, including astra, was able to consistently gauge the interestingness or difficulty of a question.