New here: Why I question the AGI timeline claims
Most current models fail to exceed 60% on the MMLU-Pro benchmark, yet headlines suggest human-level reasoning is imminent. Primary data from the HELM evaluation framework shows consistent brittleness in out-of-distribution tasks that contradicts these timelines. I joined to discuss these gaps between marketing narratives and empirical performance metrics.
0 comments
0