Aniruddha says Ox Alpha supports many agent tool calls and search-based features and that everything tested so far worked well; without tasks, traces, or a success definition, this is a tool-use experience awaiting reproduction.
Entry point: A reply to OpenCode's Ox Alpha usage post; the client configuration is not disclosed.
Tasks: No task set is published; the author mentions many agent tool calls and search capabilities.
Outcome: The author says that everything tested so far worked well.
Not disclosed: Tool list, search sources, call count, errors, latency, model version, and baseline model.
Turn this experience into a focused test of tool selection and search evidence chains. Do not interpret “many tool calls” as tool-call accuracy, or “worked well” as a benchmark score.
Fix tool schemas, search engine, permissions, and model version; prepare 20 tasks that require research before execution.
Require each run to report queries, citations, tool arguments, recovery from failures, and the final result.
Judge citation relevance, tool-argument correctness, completion rate, and rework with human labels or a test suite.
Run a fixed baseline with the same tool permissions and report model errors separately from tool/network errors.
Ox Alpha