Evals and Run Operations

What’s New

File-Backed Task Evals Task evals can replay reference runs that used direct image or PDF inputs. Stored processed values are reused, and direct model-facing files are restored from Project storage for the candidate run. See Evals.

Detailed Agent Run Operations Agent run views now combine total cost, status, failure reasons, phase timing, tool and skill activity, approval duration, actor identity, and the connection account used by each integration call. Agent lists also support server-side name and creator filters with grid and table layouts.

Azure OpenAI and Claude Fable 5 Tasks can use supported Azure OpenAI deployments, and Claude Fable 5 is available for demanding reasoning and agentic workloads.

Improvements

  • Leaving an agent page no longer cancels its live run; reloads warn before disconnecting an active stream
  • Deleting an integration preserves historical revision references and reports visible blocking dependencies
  • Approval resumes preserve the original run actor identity
  • Task fallback changes discard incompatible stale temperature settings
  • Agent phase timings remain monotonic and consistent