Trace-Bench
A benchmark and RL training substrate for evaluating causal diagnostic reasoning in AI agents TL;DR Ad-Agentic-Bench is a benchmark of 6,000 synthetic episodes that evaluates an AI agent’s abili...
A benchmark and RL training substrate for evaluating causal diagnostic reasoning in AI agents TL;DR Ad-Agentic-Bench is a benchmark of 6,000 synthetic episodes that evaluates an AI agent’s abili...
A Framework for Understanding RL with Verifiable Signals 1. The Fundamental Object: Mapping Complexity RL with verifiable rewards asks a model to learn a function: [f: \mathcal{S} \rightarrow ...
Adaptive context management for LLMs via Primal-Dual Thompson Sampling TL;DR LLMs have finite context windows, and current methods for deciding what to keep (recency truncation, summarization) a...
A simulation environment and LLM benchmark for search advertising bidding strategies TL;DR Ad Arena provides a simulation environment for search advertising auctions, extended into a public LLM ...