This is one of my all time favorite papers:
openreview.net/forum?id=ByJ...
It shows that, under fair experimental evaluation, lstms do just as well as a bunch of “improvements”
openreview.net
On the State of the Art of Evaluation in Neural Language Models
Show that LSTMs are as good or better than recent innovations for LM and that model evaluation is often unreliable.