Re-evaluating ReAct: New paper shows prompt sensitivity in multi-step agents
A recent arXiv preprint demonstrates that minor variations in ReAct prompt templates cause significant performance divergence on complex planning tasks. The authors tested five open-source models across three benchmarks, finding that accuracy swings by up to 18 percent based solely on instruction phrasing. This suggests current evaluation suites may underestimate the fragility of reasoning loops in production agents.
0 comments
0