Top
Best
New

Posted by spIrr 11 hours ago

Handbook.md shows that long policy documents do not reliably govern agents(arxiv.org)
279 points | 177 commentspage 5
mblangie 6 hours ago|
[flagged]
honestpnl 9 hours ago||
[flagged]
joka88xj 10 hours ago||
[flagged]
hnea3ekp5i 9 hours ago||
[dead]
leetrout 10 hours ago|
HANDBOOK.md is a benchmark for long-context agentic instruction following, modeled on how enterprise employees follow company handbooks in their day-to-day work. Each task is a unique RL environment with internal tools and external MCP servers, spanning five enterprise domains: Finance, Medical Billing, Insurance, Logistics, and HR.

The prompts reflect the actual jobs enterprise workers perform every day. Each task drops an AI agent into a live company environment, requiring them to cross-reference an extensive, multi-section handbook against a cluttered inbox, a multi-channel Slack workspace, Jira queues, and a stack of files (spreadsheets, PDFs), and working out both what to do and what the handbook forbids.

https://github.com/surge-ai/handbook/tree/main

ghostly_s 9 hours ago|
Please don't paste walls of text into the comment field without quotation marks. It wastes all of our time.
leetrout 5 hours ago||
It was copy/paste from my phone and when I posted there was no context / other comments and the github link was buried in the footer of the PDF of the paper.