The latest work of my PhD was accepted at the 64th Annual Meeting of the Association for Computational Linguistics in San Diego, California.
Even the strongest models we tested reach near-perfect accuracy on non-garden-path structures (93.7% for GPT-5) while collapsing on garden-path ones (46.8%).