Microsoft's DELEGATE-52 benchmark shows frontier AI models lose up to 50% document fidelity in long workflows, with catastrophic corruption in
This story is only covered by news sources that have yet to be evaluated by the independent media monitoring agencies we use to assess the quality and reliability of news outlets on our platform. Learn more here.
A Microsoft Research study warns that AI agents lose reliability when they have to run long chains of tasks, with documentary degradation and critical corruption in most of the scenarios analyzed.