Lynkr published ITSMBench results obtained by calling GPT-5.6 Sol through pi: 89 enterprise IT service-desk tasks, Pass@1/Pass@2, cost, prompt-cache hit rate, and a breakdown by task family, along with a discussion of the limitations of binary scoring.
The following is the visible body text extracted during this visit. It includes page navigation, machine translation, advertising, comments, and other page elements; verify against the original link before citing it.
To view keyboard shortcuts, press question mark View keyboard shortcuts Home Explore Notifications Chat Grok Premium History Creator Studio Articles Profile More Post 潦草学者 Scholar.L @liaocaoxuezhe Post See new posts Conversations Lynkr @LynkrDev Show translation ITSMBench Results for Lynkr
We put Lynkr through ITSMBench: 89 enterprise IT service-desk tasks, hidden end-state verifiers, with pi + gpt-5.6-sol routed through Lynkr.
The results:
• 31.0% Pass@1 — full 89-task suite • 35.0% Pass@1 / 40.0% Pass@2 — matched 2-attempt methodology • ~0% routing delta vs the model's native harness (35.51%) • $0.87–$0.90/task vs $1.29 native • 92–95% prompt-cache hit rate And the interesting part: The benchmark's binary score hides how close many failures were. 16 failed tasks still completed ≥80% of verifier assertions. One missed passing by 1 assertion out of 32.
In the hardest family, offboarding, failed tasks often completed 60–90% of required actions.
So a task that is 95% correct scores the same as 0%.
By family:
IAM: 6/9 Ops: 4/5 GRC: 3/7 Incident response: ~35% BEC/compromise: 1/5 IPAM/network: 0/5 Offboarding/endpoints: 1/22
For context, the official 5-attempt leaderboard:
Opus-5: 46.07% / $1.75 Grok-4.5: 45.39% / $0.71 GPT-5.6-sol xhigh: 39.10% / $1.53 GPT-5.6-sol high: 35.51% / $1.29 → Lynkr + GPT-5.6-sol high: 31–35% / ~$0.87
The takeaway isn't that Lynkr magically makes the model smarter.
It's that the routing layer adds essentially zero performance degradation while cutting inference cost significantly. And the assertion-level results suggest there's a lot more signal in these agent benchmarks than a binary pass/fail number reveals. Full results coming soon.
@atomicwork
@badlogicgames
@LaudeInstitute
#AIAgents #LLM AI-generated 2:28 PM · August 17, 2026 · 45 View 1 Post your reply
Reply Lynkr @LynkrDev · 2 hours Repo: - GitHub - Fast-Editor/Lynkr: Streamline your workflow with Lynkr, a CLI tool that acts as an HTTP... From github.com 4 Relevant people Lynkr @LynkrDev Follow Developer | AI Enthusiast What's trending What's new Entertainment · Trending Hayden Trending: Heroes, Kairi Entertainment · Trending Remember the Titans Trending: #Lanterns, Hal Jordan Ice Princess Trending in the United States Sheryl Yoast Show more Terms · Privacy · Cookie · Accessibility · United States TIDA · Ads info · More © 2026 X Corp.
GPT-5.6 Sol