Show HN: Agentmetry – local-first flight recorder for AI coding agents
Hacker News surfaced this AI signal from agentmetry.ai: Show HN: Agentmetry – local-first flight recorder for AI coding agents.
Topic
Coding assistants, repository agents, IDE copilots, and software-building workflows.
Latest Signals
Hacker News surfaced this AI signal from agentmetry.ai: Show HN: Agentmetry – local-first flight recorder for AI coding agents.
IssueTrojanBench is a new benchmark designed to test AI coding agents' responses to malicious issue requests. It aims to evaluate the security and robustness of AI coding tools.
JetBrains has introduced the Kotlin Benchmark to evaluate AI coding agents on real-world Kotlin tasks. This benchmark aims to provide a standardized way to assess AI performance in coding Kotlin.
No summary available yet.
Hacker News surfaced this AI signal from cursor.com: Cursor for iPad.
DeepSeek V4 Flash now runs updated weights by default on AI Gateway, improving its agentic capabilities. Its Terminal-Bench score increased from 56.9 to 82.7 without any change to the model ID or user code.
Laguna S 2.1 from Poolside now supports 10 times more capacity on Vercel's AI Gateway for both paid and free versions. This upgrade enables handling significantly higher volumes of requests, benefiting agentic coding and long-running tasks.
Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification. To expand this su...
On AI Gateway , GPT-5.6 Luna and GPT-5.6 Terra are now cheaper and <a href="https://...
Scientists are using AI coding agents to modernize scientific computing, speeding up software development and research in fields like genomics. This approach enhances the efficiency of complex computational tasks.
AI coding agents can inadvertently expose sensitive credentials during supply chain attacks. Docker Sandboxes offer a solution by isolating secrets from the agent's access.
Code Finder runs its own search loop and hands your coding agent the exact files and line ranges, faster and cheaper than the agent searching on its own.
The eighth article in Microsoft's series on Agent Experience (AX) discusses how to effectively evaluate AI coding agents and improve their integration with technology. It covers controlling the agent stack, measuring extension impact, and iterating for better results.
Precursor, our new continuous behavioral validation engine for bot management, offers visibility into how humans and bots actually interact across the full user journey. By turning session-level behavior into bot detection signals, it identifies advanced auto...
This is the seventh article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can’t control in the agent stack, how to measure whether your extensi...
Hacker News surfaced this AI signal from cognition.com: Devin Desktop, Replacing Windsurf.
A longitudinal study of a mid-sized AI-forward company reveals that mandating AI coding tools led to a doubling of merged pull requests per engineer since mid-2025. The study analyzed data from 802 developers and over 196,000 pull requests.
This is the sixth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can’t control in the agent stack, how to measure whether your extension...
AI coding agents are producing more code than ever, but the world still runs on massive, decades-old codebases. Why owning and understanding them may be the hardest job in software.
<a href="https://yenlingkuo.com"...
For a quarter century, the Google search box has been one of the most recognizable interfaces in computing: a thin white rectangle, a blinking cursor, a few typed words, and a list of blue links. On Tuesday, Google will formally retire that paradigm. <p...
Data from 1,281 agent runs across 40+ large open source repos reveals five repeatable failure patterns in coding agents, and the infrastructure fixes for each.
Replit gives professionals a secure place to build with AI. Replit Agent already protects your apps as you build by automatically scanning for vulnerabilities, and audits dependencies, before your projects are ever published. Before coding agents, a full pre-...
The initial findings from CodeScaleBench, a new benchmark designed to evaluate coding agents against the true complexity of enterprise software development, including large codebases and multi-repository tasks.
Today, we're announcing Sourcegraph 7.0, a release that marks the beginning of a new chapter for our company and product.
Coding assistants, repository agents, IDE copilots, and software-building workflows.