Models, labs and research · Straight from the source

The Tech Gazette

AI Edition
Vol. I No. 2 ★★★ Wednesday, September 23, 2026 ★★★ Closing 11:03 PM BRT

News

Announcements from the labs, companies and platforms themselves.

Today

  1. Meta Engineering9:00 PMengineering.fb.com
    Bringing Private Processing to Meta AI Glasses

    We believe glasses are the best form factor for having AI help throughout your day. They can understand your personal context better than other kinds of devices and keep you present without picking up a mobile phone. Most of the time, glasses are helping you see well, protecting…

    Other
  2. Vercel9:00 PMvercel.com
    Vercel Connect now supports TanStack AI

    Agents built with TanStack AI can now call OAuth-protected MCP servers through Vercel Connect, with no credentials for you to store or rotate. The new @vercel/connect/tanstack-ai subpath exports connectMCPTransport, which takes a TanStack transport config and attaches a Connect-b…

    LaunchDev Tools
  3. NVIDIA Developer7:54 PMdeveloper.nvidia.com
    Introducing NV-Reason-CT Open 3D CT VLM for Radiologist Chain-of-Thought Reasoning

    Radiology AI has made remarkable strides in detecting abnormalities across chest X-rays, pathology slides, and 2D scans. Yet one of the most clinically rich and...

    LaunchModels
  4. GitHub Changelog6:25 PMgithub.blog
    More ways to request and configure Copilot code reviews

    GitHub Copilot code review now offers additional personal configurations to an expanded set of Copilot plans and an enterprise-level default setting. These improvements are now generally available: A dedicated personal… The post More ways to request and configure Copilot co…

    LaunchDev Tools
  5. Microsoft Research1:01 PMmicrosoft.com
    Offloaded inference for real-world physical AI robotics

    Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for r…

    Research
  6. Google DeepMind1:00 PMdeepmind.google
    Advancing Private AI Compute with secure, server-side memory

    Introducing private, server-side memory to Private AI Compute for personal AI.

    LaunchSecurity
  7. Google DeepMind12:25 PMdeepmind.google
    Gemini 3.8 text-to-speech says hello

    Also reported by Vercel

    More on Gemini 3.8 · Gemini

    Models
  8. GitHub Changelog12:00 PMgithub.blog
    Local sandboxing in the GitHub Copilot app

    Local sandboxing helps reduce the potential impact of unintended commands by limiting access to files, network resources, and credentials on your machine. In the GitHub Copilot app, you configure it… The post Local sandboxing in the GitHub Copilot app appeared first on The…

    Security
  9. JetBrains10:23 AMblog.jetbrains.com
    Small Talk With Prasun Kumar, CEO and Founder of Oppex AI

    What is Small Talk? Small Talk is our short Q&A with founders from the JetBrains Startup Program. They answer a handful of questions in their own words about what they’re building, why they started, and what they’ve figured out along the way. No pitch, no polish –…

    Dev Tools
  10. OpenAI9:00 AMopenai.com
    Sam Altman’s remarks at the United Nations Security Council

    OpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council.

    Other
  11. Anthropic9:00 AManthropic.com
    Claude discovers a novel enzyme system

    In early results from our new life sciences research lab, Claude agents found an enzyme system whose function is still unknown.

    Research
  12. Pinecone9:00 AMpinecone.io
    Pinecone BYOC: Trusted AI Knowledge in the Customer Cloud

    Pinecone Bring Your Own Cloud is generally available. The data plane runs in your cloud account, and Pinecone manages it without inbound access.

    LaunchPlatforms
  13. OpenAI7:00 AMopenai.com
    Introducing MentalHealthBench

    MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.

    LaunchResearch

Yesterday

  1. NVIDIA11:30 PMblogs.nvidia.com
    At AI Day Singapore, NVIDIA and Partners Showcase AI Advancements Across Southeast Asia

    NVIDIA AI Day Singapore, which takes place Sept. 22-23 at the Raffles City Convention Centre, is offering attendees opportunities to explore the hands-on training, expert-led sessions and advanced tools to accelerate their work in AI and high-performance computing. At the event,…

    Other
  2. GitHub Changelog11:14 PMgithub.blog
    OpenTelemetry in the GitHub Copilot app

    Understand how Copilot agents perform and interact with models and tools. The GitHub Copilot app now supports OpenTelemetry (OTel) configuration through enterprise-managed settings. OTel is an open source observability framework.… The post OpenTelemetry in the GitHub Copilo…

    LaunchDev Tools
  3. GitHub Changelog9:34 PMgithub.blog
    New features and improvements in Copilot for JetBrains

    GitHub Copilot for JetBrains 1.18.0 brings AI-assisted tool approvals, more control over agent conversations, and shared skills and instructions for your organization. You can also review plans with the Codex… The post New features and improvements in Copilot for JetBrains…

    LaunchDev Tools
  4. What's New with AWS9:06 PMaws.amazon.com
    Amazon CloudWatch Omni: AI-first observability for agents and applications

    AWS announces the general availability of Amazon CloudWatch Omni, an evolution of Amazon CloudWatch. Omni is an AI-powered observability experience organized around your teams and the applications they run, so that you can observe and troubleshoot your applications and agents in…

    LaunchPlatforms
  5. Fireworks AI9:00 PMfireworks.ai
    Every byte counts: ARCv3 and the case for cross-region RL

    Fireworks’ ARCv3 delivers lossless BF16 weight updates nearly 50% smaller than ARCv2, making cross-region reinforcement learning more practical.

    LaunchModels
  6. Fireworks AI9:00 PMfireworks.ai
    Introducing Ember-1

    Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens.

    More on Ember-1 · Ember

    LaunchModels
  7. Vercel9:00 PMvercel.com
    Drives for Vercel Sandbox are now in public beta

    Drives for Vercel Sandbox are now available in public beta on Hobby, Pro, and Enterprise. A Drive is persistent storage that you mount as a directory in a Vercel Sandbox. It isn’t tied to a single sandbox, so you can reuse the same Drive across runs and different sandbox instance…

    LaunchPlatforms
  8. Cursor9:00 PMcursor.com
    Rollouts and Security Review

    Two bots for the last mile of shipping code: Rollouts watches every change as it deploys, and Security Review reports exploitable bugs on every pull request.

    LaunchDev Tools
  9. GitHub Changelog7:24 PMgithub.blog
    Faster C++ code intelligence with whole codebase indexing

    C++ code intelligence in GitHub Copilot CLI is now faster with support for whole codebase indexing. C++ repositories can contain millions of lines of code across deeply connected source files… The post Faster C++ code intelligence with whole codebase indexing appeared first…

    LaunchDev Tools
  10. AWS News7:23 PMaws.amazon.com
    Introducing Amazon CloudWatch Omni: collaborative AI-powered observability for your applications

    Amazon CloudWatch Omni is the next evolution of CloudWatch — unified observability that brings your applications and AI agents into one reimagined experience, with auto-discovered topology, natural language queries, and AI-guided investigation powered by AWS DevOps Agent.

    LaunchPlatforms
  11. AWS News7:21 PMaws.amazon.com
    Introducing Amazon CloudWatch Omni: AI-powered observability for generative AI and agentic workloads

    Learn how Amazon CloudWatch Omni delivers AI-powered observability purpose-built for generative AI and agentic workloads. Trace, evaluate, and experiment with AI agents across any framework—directly from your IDE or a standalone web experience—using open standards and built-in ev…

    LaunchDev Tools
  12. OpenAI3:00 PMopenai.com
    Introducing GPT-6 Sol and Luna

    Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.

    Also reported by Vercel · Netlify · OpenAI API Changelog · GitHub Changelog · What's New with AWS · AWS Machine Learning

    More on GPT-6 · GPT

    LaunchModels
  13. Databricks1:00 PMdatabricks.com
    Genie One MCP: Give any AI Agent the Right Business Context

    Business leaders often have access to plenty of data, but still can’t get a reliable...

    Dev Tools
  14. Databricks1:00 PMdatabricks.com
    The Genie One MCP is now Generally Available

    AI coworkers and coding agents are spreading fast across organizations, and each...

    LaunchDev Tools
  15. DigitalOcean12:01 PMdigitalocean.com
    Introducing DigitalOcean Managed Agents: One AI-native stack to power your intelligence

    Following a successful private preview, we’re thrilled to open DigitalOcean Managed Agents to everyone. Teams can now deploy their preferred agent harness (like OpenCode, Codex CLI) or bring their own, connect agents to 16,000+ tools, and help agents do more work at scale without…

    LaunchDev Tools
  16. Databricks10:57 AMdatabricks.com
    Unity Catalog Pages: a governed home for your business knowledge in Genie Ontology

    Every AI agent is only as good as the context it is grounded in. Ask an agent a question...

    Platforms
  17. NVIDIA9:00 AMblogs.nvidia.com
    NVIDIA Isaac ROS 5.0 Advances Agentic, Open Source Robotics Development

    To build and deploy sophisticated robotics applications that can perceive, reason and act in dynamic environments, developers need new physical AI models and tools. The ROS open framework is a project from Open Robotics that helps humans build robots. NVIDIA Isaac ROS 5.0 — a col…

    More on ROS 5.0 · ROS

    LaunchDev Tools
  18. Claude API Release Notes9:00 AMplatform.claude.com
    Claude Opus 5.5

    We've launched Claude Opus 5.5 (claude-opus-5-5), a model for long-running agentic coding and knowledge work. It has a 1M token context window by default, 128k max output tokens, and always-on adaptive thinking, at $4 / $20 USD per MTok (Claude Opus 5 is $5 / $25). Claude Opus 5.…

    Also reported by Vercel · Netlify · What's New with AWS · What's New with AWS · Microsoft Azure · GitHub Changelog · AWS Machine Learning

    More on Opus 5.5 · Opus

    LaunchModels
  19. Gemini API Changelog9:00 AMai.google.dev
    Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS generally available (GA)

    Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS generally available (GA): Released our next-generation text-to-speech (TTS) audio models and the Gemini API Voices endpoint (/v1beta/voices): Gemini 3.8 Flash TTS (gemini-3.8-flash-tts): Flagship creative TTS model engineered for…

    More on Gemini 3.8 · Gemini

    LaunchModels
  20. JetBrains6:00 AMblog.jetbrains.com
    JetBrains Air: Building a System of Products for Agentic Software Development

    AI can produce code. Organizations still have to produce software. Agentic development is changing how software gets made, but it hasn’t changed what it costs to be wrong. Six months ago, we began publicly experimenting with agentic development environments. Around the same time,…

    LaunchDev Tools
  21. Google Cloud release notes4:00 AMdocs.cloud.google.com
    September 22, 2026

    BigQuery Feature You can now publish a BigQuery data agent in Gemini Enterprise by registering the agent with Agent Registry and importing it using default Google-managed credentials. When BigQuery and Gemini Enterprise are in the same Google Cloud project and configured with a m…

    LaunchPlatforms
  22. What's New with AWS1:00 AMaws.amazon.com
    Amazon Connect Customer launches agent-to-agent collaboration

    Amazon Connect Customer now supports agent-to-agent collaboration, giving customers the choice to bring in specialized AI agents during a live interaction to resolve a customer request. With this launch, Connect Customer AI agents can collaborate with each other, and with AI agen…

    LaunchPlatforms

Monday, September 21

  1. OpenAI9:00 PMopenai.com
    Priorities and principles for effective third party assessments

    OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.

    Other
  2. Hugging Face9:00 PMhuggingface.co
    Transformers now runs llama.cpp quants
    LaunchModels
  3. Hugging Face9:00 PMhuggingface.co
    Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community
    Other
  4. Fireworks AI9:00 PMfireworks.ai
    Introducing The Specialized Intelligence Index

    Fireworks introduces the Specialized Intelligence Index, a one-stop destination for real-work benchmarks across industries.

    LaunchResearch
  5. Upstash9:00 PMupstash.com
    Context7 Search: A Grounding API for Coding Agents

    Ground your coding agent in official documentation with one REST request. Includes a Vercel AI SDK example that uses Context7 Search as a tool.

    LaunchDev Tools
  6. NVIDIA Developer6:51 PMdeveloper.nvidia.com
    Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton

    The compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability...

    LaunchDev Tools
  7. Amazon Science4:31 PMamazon.science
    Amazon launches research initiative with Stanford University to advance AI and science

    The collaboration aims to advance research while broadening participation and translating discovery into real-world solutions.

    Research
  8. AWS Machine Learning3:30 PMaws.amazon.com
    xAI’s Grok 4.6 is now available in Amazon Bedrock

    xAI's Grok 4.6 is now available in Amazon Bedrock: a frontier model for long-running agents, coding, and knowledge work, with a 500K token context window and four reasoning effort levels. It runs on both the bedrock-mantle and bedrock-runtime endpoints, with Converse API and cros…

    More on Grok 4.6 · Grok

    LaunchModels
  9. Amazon Science2:51 PMamazon.science
    Advancing AI for biology: Teaching models to design and characterize antibodies

    Three new papers from Amazon Bio Discovery address bottlenecks in AI-driven antibody engineering, from benchmarking binding predictors to experimentally validating de novo design.

    Research
  10. LangChain2:14 PMlangchain.com
    Jev is now available in LangSmith Evals

    Use Jev as a judge for LangSmith evals to evaluate agent traces with faster, cheaper structured feedback across production runs, datasets, and regression tests.

    LaunchDev Tools
  11. Google Cloud1:00 PMcloud.google.com
    Scale your AI workloads faster and more efficiently with GKE Pod snapshots

    When running modern AI workloads, there’s often a conflict between performance and cost. Workloads like large language models (LLMs) load massive files, and may serve thousands of AI agents that need to execute code instantly. If each component is starting “cold” with a full data…

    Platforms
  12. Microsoft Research12:30 PMmicrosoft.com
    Improving synthesis prediction of small molecules at scale with RetroChimera

    Custom-made molecules are advancing medicine, materials, and agriculture, but producing them is slow and expensive. A new Nature paper highlights RetroChimera, a predictive model that helps accelerate chemical synthesis, helping researchers explore a wide range of molecules. The…

    LaunchModels
  13. NVIDIA11:51 AMblogs.nvidia.com
    AI Security Is an Engineering Problem — How to Solve It at Every Layer of the Agent Stack

    AI security is an engineering problem. That means defined security requirements, enforceable controls, named owners and evidence that protections work. As AI becomes more capable, the industry must accelerate security engineering, broaden access to defensive tools and share what…

    Security
  14. OpenAI9:00 AMopenai.com
    Advisory Group on Mathematics and Artificial Intelligence

    OpenAI is working with an independent Advisory Group on Mathematics and Artificial Intelligence to guide the review and communication of emerging AI results.

    Other
  15. OpenAI9:00 AMopenai.com
    Higgsfield AI ships new video features in a day with GPT-6 Astra

    With GPT-6 Astra, Higgsfield AI makes video ad creation easier for small businesses and brings new creative tools to market faster.

    Also reported by Microsoft Azure · OpenAI · OpenAI

    More on GPT-6 · GPT

    LaunchOther
  16. OpenAI7:00 AMopenai.com
    Building standards for the next phase of AI

    OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety.

    Other
  17. OpenAI4:00 AMopenai.com
    Expanding OpenAI Academy with new learning paths

    Explore new OpenAI Academy learning paths for employees, developers, leaders, educators, and students to build and demonstrate practical AI skills.

    Other

Sunday, September 20

  1. Fireworks AI9:00 PMfireworks.ai
    The frontier isn’t a model. It’s a router.

    18 models, 113 coding tasks. The best single model gets 74.1% at $6.52. A perfect router gets 97.6% at $1.88. See how FireRouter closes the gap.

    Models
  2. Vercel9:00 PMvercel.com
    MiMo V2.6 models now available on AI Gateway

    MiMo V2.6 Pro, MiMo V2.6 Flash, and MiMo V2.6 Pro UltraSpeed from Xiaomi are now available on AI Gateway. MiMo V2.6 combines coding, reasoning, and tool use with native text, image, audio, and video understanding. Its 1M token context supports long repositories, tool traces, and…

    LaunchModels
  3. Vercel9:00 PMvercel.com
    AI Gateway now supports TypeSafe clients and an HTTP API for Jev

    You can now call Jev from TypeSafe AI through AI Gateway using an existing TypeSafe client or the HTTP API, in addition to the AI SDK. TypeSafe client: Point an existing TypeSafe client at AI Gateway without changing its evaluation calls. HTTP API: Call Jev directly from any lang…

    LaunchDev Tools
  4. Vercel9:00 PMvercel.com
    Grok 4.7 now available and 40% off on AI Gateway, fx, and eve

    Grok 4.7 from SpaceXAI is now available on AI Gateway and 40% off through September 27. The discount applies automatically when you call spacexai/grok-4.7. Grok 4.7 has a 500K token context window and supports low, medium, high, and xhigh reasoning levels, giving you control over…

    More on Grok 4.7 · Grok

    LaunchModels
  5. Sourcegraph9:00 PMsourcegraph.com
    The autonomous codebase

    What's left for us to build?

    Dev Tools

Friday, September 18

  1. Vercel3:00 PMvercel.com
    WebMCP support now available in mcp-handler

    mcp-handler now has experimental support for WebMCP, the proposed web standard for exposing tools to in-browser agents. Add a single script tag to your site, and your existing MCP tools become available there too. Opt tools in by adding them to the experimental_webMcp object: The…

    LaunchDev Tools
  2. AWS Machine Learning1:52 PMaws.amazon.com
    Introducing Kimi K3 on Amazon Bedrock

    Kimi K3 from Moonshot AI is now available on Amazon Bedrock, giving you a powerful new open-weight option for coding and knowledge work. It offers native vision, a 1-million-token context window, and explicit prompt caching to reduce latency and input costs.

    LaunchModels
  3. Google Cloud1:30 PMcloud.google.com
    Announcing Native BM25 Ranking in AlloyDB and Cloud SQL

    Vector search is a critical component of generative AI, retrieval-augmented generation (RAG), and data agent architectures, but sometimes vector search alone isn't enough. While vector embeddings are incredible at understanding conceptual meaning, they stumble on specific alphanu…

    LaunchPlatforms
  4. Google Cloud1:00 PMcloud.google.com
    Changing the game: Using agentic AI to secure infrastructure code

    AI is accelerating software development at an unprecedented pace. But as code generation scales, so do the challenges of securing the code, especially emerging AI-based vulnerability exploitations. To meet these challenges, the Google AI and Infrastructure team is transforming ho…

    Security
  5. AWS Machine Learning12:31 PMaws.amazon.com
    The new AgentCore runtime: Elastic, optimized, and consistently fast starts

    Today we are announcing the new AgentCore runtime, a capability of Amazon Bedrock AgentCore built for the speed, flexibility, and cost efficiency that production agents demand. It reclaims memory as sessions release it and delivers consistent cold starts regardless of image size…

    LaunchPlatforms
  6. What's New with AWS12:15 PMaws.amazon.com
    Kimi K3 by Moonshot AI is now generally available on Amazon Bedrock

    Amazon Bedrock continues to expand its open weight model portfolio with the same security and governance that customers rely on. Today, Kimi K3 from Moonshot AI is generally available on Amazon Bedrock, giving you a powerful new option for coding and knowledge work. Accordin…

    LaunchModels
  7. The GitHub Blog12:00 PMgithub.blog
    Should you read the code, is RAG dead, and did Skills kill MCP?

    We dive into these questions and other AI hot takes on the latest episode of the GitHub Podcast. The post Should you read the code, is RAG dead, and did Skills kill MCP? appeared first on The GitHub Blog.

    Dev Tools
  8. JetBrains10:52 AMblog.jetbrains.com
    The AIDEs Framework: How We Built a “Theory of Everything” for AI Development Tools

    We are living in genuinely interesting times. AI is disrupting software development at a pace where new models, tools, and practices appear almost daily. Many teams’ natural first instinct is to spend ever more time chasing updates. After almost two years of AI product and market…

    Dev Tools
  9. JetBrains10:42 AMblog.jetbrains.com
    Making Local AI Smarter and Faster

    We want local coding agents to be smart and fast, with the ability to understand a codebase, do useful work, and finish tasks without long waits. This Junie Local update makes it practical to use a more capable model on your own machine. In the first release, we had to choose bet…

    LaunchModels
  10. What's New with AWS10:25 AMaws.amazon.com
    The new AgentCore Runtime is now available in Amazon Bedrock AgentCore

    Today, AWS announces the availability of the next generation of AgentCore Runtime, the serverless microVM compute within Amazon Bedrock AgentCore. The new Runtime delivers elastic memory management that reclaims unused memory throughout the session so you pay for actual usag…

    LaunchPlatforms
  11. Anthropic9:00 AManthropic.com
    Partnering with Accenture on embedded evaluation

    We’re partnering with Accenture on independent evaluation of frontier AI—part of our recent commitment to embed evaluators at Anthropic. Both we and Accenture expect to invest at least $1 billion to build capacity in this area over the next five years.

    Other
  12. Claude API Release Notes9:00 AMplatform.claude.com
    The Compliance API local session endpoints now also return transcripts of Claude in Chrome sessions (product_surface value…

    The Compliance API local session endpoints now also return transcripts of Claude in Chrome sessions (product_surface value claude_in_chrome), in beta for Claude Enterprise organizations, with your existing Compliance Access Key and the read:compliance_user_data scope. See Session…

    LaunchPlatforms
  13. Gemini API Changelog9:00 AMai.google.dev
    Gemini 2.5 models access update

    Gemini 2.5 models access update: To ensure reliable performance for everyone, we are limiting access to the 2.5 models to users who have actively used them in the past. These models are not deprecated and will continue to be served until further notice through the API. For any ne…

    More on Gemini 2.5 · Gemini

    Models
  14. JetBrains8:51 AMblog.jetbrains.com
    Your diff has a demo now with Junie /demo

    You have the change ready and the tests are green. Now someone has to launch the app, find the right screen, and check the flow. Often, that someone is still you, even when an agent helped write the code. You should be able to delegate that part too. Junie /demo is a new mode in…

    LaunchDev Tools
  15. Vercel4:00 AMvercel.com
    Jev is the fastest-adopted model in AI Gateway history

    Within 24 hours of launching on AI Gateway, Jev from TypeSafe AI reached more than twice as many paid teams as any previous model launch, making it the fastest-adopted model in gateway history. Jev passed every other comparison model in its first twelve hours and continued to wid…

    LaunchModels
  16. Google Cloud release notes4:00 AMdocs.cloud.google.com
    September 18, 2026

    Agent Platform Workbench Change 20260918-2230-rc0 Release Change 20260918-2230-rc0 Release Change Installed latest packages from upstream dependencies. Change Installed latest packages from upstream dependencies. Feature JupyterLab now forwards client-side logs (console errors, u…

    LaunchDev Tools

Thursday, September 17

  1. Apple Machine Learning9:00 PMmachinelearning.apple.com
    Dynamically Scaled Activation Steering

    Activation steering has emerged as a powerful method for guiding the behavior of generative models towards desired outcomes such as toxicity mitigation. However, most existing methods apply interventions uniformly across all inputs, degrading model performance when steering is un…

    Research
  2. Vercel9:00 PMvercel.com
    GLM 5.3 FlashX now available on AI Gateway

    GLM 5.3 FlashX is now available on AI Gateway. GLM 5.3 FlashX is a high-speed serving option for Z.ai's multimodal coding model, delivering inference at ~200 tokens per second for faster streamed responses. The higher serving speed is useful for coding agents, tool loops, and int…

    More on GLM 5.3 · GLM

    LaunchModels
  3. Google Research5:45 PMresearch.google
    The future of practice: Enabling teachers to create learning interactives with generative UI

    Education Innovation

    Research
  4. Vercel4:00 PMvercel.com
    Run Terminal-Bench and other Harbor evals on Vercel Sandbox

    You can now run Harbor evals on Vercel Sandbox. Harbor is the open-source harness behind Terminal-Bench, whose registry includes many other benchmarks such as SWE-bench, tau3-bench and OSWorld. Pass --env vercel to harbor run and each trial executes in its own isolated Firecracke…

    LaunchPlatforms
  5. Vercel3:00 PMvercel.com
    The skills CLI now supports Notion hosted skills

    skills@1.7.0 adds Notion skills databases as an install source for agent skills. Notion skills are reusable agent skills written as Notion pages. Teams author, review, and update them in the workspace they already use, then install them into any agent the skills CLI supports. No…

    LaunchDev Tools
  6. LangChain2:40 PMlangchain.com
    Building an Agent Harness for Life Sciences: Introducing Deep Life Sci

    Deep Life Sci is LangChain's open source agentic assistant for clinical and lab scientists. It pulls from 600K+ ClinicalTrials.gov studies, 29M PubMed abstracts, and 12M PubMed Central full-text articles, with sandboxed sub-agents for real data analysis.

    LaunchDev Tools
  7. Databricks2:00 PMdatabricks.com
    The Web Search Your Agent Inherited Isn't Good Enough

    An agent that needs the outside worldAn engineer at a software company is building...

    Dev Tools
  8. Lambda1:29 PMlambda.ai
    Closing the loop: agentic evaluation for image editing foundation models

    Why evaluating image editing models is both critical and challenging Instruction-based image editing is becoming a core capability of multimodal foundation models. Users can increasingly edit images simply by describing what they want: “remove the person in the background,” “make…

    Research
  9. Pinecone12:12 PMpinecone.io
    VQ-bench: a Composable Vector Quantization Framework

    Most published quantizers are built from the same small set of primitives. VQ-bench is an open-source library of those primitives, plus a reproducible benchmark of 14 quantizers across VIBE datasets.

    Research
  10. Pinterest Engineering12:01 PMmedium.com
    Beyond Two Towers: Launching the 3-Tower Engagement Co-Train Model (Part 2)

    Authors: Longyu Zhao (Staff Machine Learning Engineer), Gwendolyn Zhao (Staff Machine Learning Engineer), Peng Yan (Senior Machine Learning Engineer), Yuanlu Bai (Senior Machine Learning Engineer), Yuan Wang (Senior Machine Learning Engineer), Yao Cheng (Staff Machine Learning En…

    Models
  11. Android Developers11:06 AMandroid-developers.googleblog.com
    Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks

    Posted by Matthew McCullough, VP, Product Management, Android Developer When we first launched Android Bench, we built a rigorous foundation for evaluating how large language models (LLMs) assist developers with real-world Android tasks. As AI models and agents rapidly evolve, we…

    LaunchResearch
  12. Anthropic9:00 AManthropic.com
    Introducing the Life Sciences Verification Program

    Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

    LaunchOther
  13. Gemini API Changelog9:00 AMai.google.dev
    Antigravity Agent 09-2026

    Antigravity Agent 09-2026: Released antigravity-preview-09-2026, which replaces and deprecates antigravity-preview-05-2026. If you run on a remote sandbox (environment: "remote") and read only output_text or model_output steps, update the agent string and nothing else changes. If…

    LaunchDev Tools
  14. Neon9:00 AMneon.com
    The Neon backend is GA: a complete set of primitives so agents can build

    Neon is now a complete suite of backend primitives built around the database and rooted on the lakebase architecture: Lakebase Postgres, Object Storage, Functions, Managed Better Auth, and AI Gateway. All tools are GA and ready for production. Tell your agent to deploy them.

    LaunchPlatforms
  15. Ai25:00 AMallenai.org
    What a crowdsourced game revealed about steering Olmo 3

    A crowdsourced game built on Olmo 3 showed how people can exploit unexpected model behaviors to stress-test prosocial AI evaluations—and how open access to a model’s internals can help researchers understand why those tests break.

    More on Olmo 3 · Olmo

    Research
  16. Vercel4:00 AMvercel.com
    Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend

    AI Gateway Production Index — September 2026 Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs. That traffic gives us a view of what AI usage actually looks like in today's enterprise, and we publish it here monthly. See the Pr…

    More on Fable 5.1 · Fable

    Models

Wednesday, September 16

  1. What's New with AWS11:37 PMaws.amazon.com
    Amazon Quick adds Generate Sheet and image-based analysis creation

    Amazon Quick now expands Generate Analysis with two new ways to create dashboards faster. You can generate a single sheet inside an existing analysis by describing it in natural language, and you can generate a new analysis from an image of an existing dashboard.    &nb…

    LaunchPlatforms

Sunday, July 26

  1. Berkeley AI Research6:00 AMbair.berkeley.edu
    Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

    .abbel-fig { display: block; text-align: center; margin: 2.4em 0; line-height: 1.4; max-width: 100%; }.abbel-fig img { display: block; margin: 0.65em auto 0; height: auto; max-width: 100%; } /* Image sizes; captions use a narrower measure below */.abbel-fig--wide img { width: 100…

    Research

New Models

Everything that joined the OpenRouter catalog.

  1. OpenRouterSep 23
    New model: inclusionAI: Ming Image 0.1 Design Layer
  2. OpenRouterSep 23
    New model: Google: Gemini 3.8 Flash Lite TTS
  3. OpenRouterSep 23
    New model: Google: Gemini 3.8 Flash TTS
  4. OpenRouterSep 23
    New model: Z.ai: GLM 5.3 Prime
  5. OpenRouterSep 23
    New model: Qwen: Qwen3.8 Max Prime
  6. OpenRouterSep 23
    New model: Recraft: Recraft V4.1 Flash
  7. OpenRouterSep 23
    New model: Space Bunny Alpha
  8. OpenRouterSep 23
    New model: AionLabs: Aion 3.5 Mini
  9. OpenRouterSep 23
    New model: AionLabs: Aion 3.5
  10. OpenRouterSep 23
    New model: Upstage: Solar Mini 4
  11. OpenRouterSep 22
    New model: Cohere: Command A+
  12. OpenRouterSep 22
    New model: OpenAI: GPT-6 Luna Pro
  13. OpenRouterSep 22
    New model: OpenAI: GPT-6 Luna
  14. OpenRouterSep 22
    New model: OpenAI: GPT-6 Sol Pro
  15. OpenRouterSep 22
    New model: OpenAI: GPT-6 Sol
  16. OpenRouterSep 22
    New model: inclusionAI: Ming Image 0.1 Design
  17. OpenRouterSep 22
    New model: Anthropic: Claude Opus 5.5
  18. OpenRouterSep 22
    New model: AssemblyAI: Universal-3.5 Pro
  19. OpenRouterSep 21
    New model: Xiaomi: MiMo-V2.6-Pro-UltraSpeed
  20. OpenRouterSep 21
    New model: Xiaomi: MiMo-V2.6-Flash
  21. OpenRouterSep 21
    New model: Xiaomi: MiMo-V2.6-Pro
  22. OpenRouterSep 21
    New model: SpaceXAI: Grok 4.7
  23. Hugging Face HubSep 21
    Open model: XiaomiMiMo/MiMo-V2.6-Flash-RL
  24. Hugging Face HubSep 21
    Open model: XiaomiMiMo/MiMo-V2.6-Pro-RL
  25. OpenRouterSep 21
    New model: Qwen: Qwen3.8 Omni Flash
  26. Hugging Face HubSep 19
    Open model: convaiinnovations/laya-multilingual
  27. OpenRouterSep 18
    New model: PrismML: Ternary Bonsai 2 27B
  28. OpenRouterSep 18
    New model: Z.ai: GLM 5.3 FlashX
  29. Hugging Face HubSep 18
    Open model: convaiinnovations/laya
  30. OpenRouterSep 17
    New model: TypeSafe: Jev Latest
  31. OpenRouterSep 17
    New model: TypeSafe: Jev 1.13
  32. OpenRouterSep 17
    New model: Pareto
  33. Hugging Face HubSep 17
    Open model: inclusionAI/Ming-Image-0.1-Design
  34. Hugging Face HubSep 16
    Open model: XingChen-AGI/Xing4.0-29B-A4B
  35. Hugging Face HubSep 16
    Open model: Cactus-Compute/needle3
  36. Hugging Face HubSep 14
    Open model: Qwen/Qwen-Image-2.1
  37. Hugging Face HubSep 12
    Open model: yandex/AliceAI-Foundation-80B-A3B-Base
  38. Hugging Face HubSep 9
    Open model: deepseek-ai/DeepSeek-V4.1-Flash

Research

New arXiv papers in AI, language, machine learning and software engineering. The 30 most relevant, as ranked by Jev.

  1. arXiv cs.CLSep 23Open code
    When Context Misleads: In-context Learning with Jurisdiction in Large Language Models

    In-Context Learning (ICL) has become a cornerstone of modern LLM deployment. However, existing ICL post-training methods have a critical blind spot: they excel at extracting patterns from demonstrations while often neglecting context authority, the ability to determine whether co…

  2. arXiv cs.AISep 23Open code
    StudentBench: AI and human tutoring yield equivalent GRE learning gains

    Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-sc…

  3. arXiv cs.CLSep 23Open code
    Exact Feedback Is Not Control: Evaluating Text-based Closed-Loop Revision in LLMs

    Closed-loop revision is increasingly used in large language model (LLM) applications, but failures may reflect incomplete feedback or ineffective responses to correct feedback. We introduce a fixed-budget revision protocol with deterministic verifiers that report all remaining vi…

  4. arXiv cs.SDSep 23Open code
    Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

    Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language models, enabling contextual understanding of auditory events. Making these capabilities practical on devices with limited memory and computation motivates our focus on small ALMs wit…

  5. arXiv cs.CLSep 23Open code
    SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

    Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather than internalized as reusable model capabilities. We introduce \texttt{SkillGym}, a framework that transforms these skills into e…

  6. arXiv cs.CLSep 23Open code
    Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

    Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo} variant reduced the decoder from 32 to 4 layers, wh…

  7. arXiv cs.CLSep 23
    Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings

    Meaning identity (whether two sentences say the same thing after wording changes) is treated in retrieval and RAG as a geometric fact about independently encoded sentence vectors. We show that, for frozen off-the-shelf encoders and language models, it is not: identity is computed…

  8. arXiv cs.AISep 23
    PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

    As Large Language Models (LLMs) evolve into autonomous agents that alter real-world states, ensuring operational safety across multi-step workflows has become a critical challenge. While recent work has moved beyond single-turn evaluation toward multi-turn paradigms, key limitati…

  9. arXiv cs.SESep 23Open code
    FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration

    Deployment requires an agent to turn application code into a running system whose services connect, become ready, and remain observable. FDE-Bench evaluates this capability with 136 deployment-configuration tasks spanning Docker images, multi-service Compose stacks, and Kubernete…

  10. arXiv cs.CLSep 23Open code
    Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

    Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine langu…

  11. arXiv cs.CVSep 23Open code
    NV-Reason-CT: 3D Visual Language Model for CT Analysis

    We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D c…

  12. arXiv cs.CLSep 23Open code
    ThaiTrees: Thai Syntactic Dependency Trees Across Domains

    Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank for training and evaluating parsers, but lacks a large automatically parsed corpus…

  13. arXiv cs.SESep 23
    Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

    Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasonin…

  14. arXiv cs.AISep 23
    An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

    Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical eviden…

  15. arXiv cs.LGSep 23Open code
    hyperbolix: Hyperbolic Deep Learning in JAX

    We present hyperbolix, an open-source library for hyperbolic deep learning in JAX, built on Flax NNX. To our knowledge, it is the first comprehensive, general-purpose hyperbolic deep learning library in JAX. It includes six manifolds with a common interface: Euclidean space, the…

  16. arXiv cs.LGSep 23Open code
    TNLearn: An Open Source Python Package for Task-based Neurons

    The brain does not rely on a single type of neuron to perform all kinds of tasks; instead, it designs different neurons for different tasks. The concept of task-based neurons represents a paradigm shift compared to task-based architectures. It argues that solving a specific probl…

  17. arXiv cs.CLSep 23Open code
    Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing

    We investigate whether state-of-the-art large language models (LLMs) generate feedback that reflects the pedagogical practices of expert teachers in terms of feedback focus and adaptivity. Previous evaluation efforts have examined feedback characteristics, its impact on learning,…

  18. arXiv cs.LGSep 23Open code
    MENO: Memory-Efficient Neural Operator

    We propose the Memory-Efficient Neural Operator (MENO) as a high-performance PDE neural solver based on the Manifold Function Encoder (MFE). MENO features three primary advantages: (1) MENO has a significantly smaller memory footprint and much faster training speed than other pop…

  19. arXiv cs.CVSep 23
    AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios

    Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning. We introduce AnchorReasoning, a visually grounded reas…

  20. arXiv cs.CLSep 23Open code
    Towards Efficient Reasoning: Learning Causal Shortcuts for Diffusion Language Models

    Diffusion Language Models (DLMs) have attracted significant attention for their strong reasoning ability. However, under a bidirectional attention mechanism, DLMs operate over an exponentially large exploration space compared to autoregressive models (ARMs), making it challenging…

  21. arXiv cs.LGSep 23Open code
    Learning Collective Dynamics with Differentiable Gaussian Representations

    Collective responses depend on individual differences, contact opportunities, and accumulated experience. Learning their dynamics from aggregate counts requires connecting a population's response distribution to both current observations and future behavior. We introduce Differen…

  22. arXiv cs.CVSep 23
    InGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation

    Modern text-to-image (T2I) models generate high-quality images from arbitrary user prompts, yet they can just as easily produce not-safe-for-work (NSFW) content. Conventional outer guardrails consist of two components: a prompt classifier that checks for risk before generation, a…

  23. arXiv cs.CVSep 23
    RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models

    Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. However, quantization affects different layer types in inconsistent ways, so identifyi…

  24. arXiv cs.CLSep 23
    Evaluating Open-Weight LLMs for Turkish Domain Documents Under Retrieval and Hardware Constraints

    Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B models for Turkish document question answering under a resource-constrained local…

  25. arXiv eess.SYSep 23Open code
    EvEMTBench: An Open Benchmark for Machine Learning in Power System Protection

    Studies of machine-learning-based power system protection are difficult to compare because task definitions, measurement access, data partitions, metrics, and generalization conditions often differ. EvEMTBench addresses this gap with an open, executable, and versioned benchmark t…

  26. arXiv cs.AISep 23
    Not What You Meant: Can LLMs Follow a Specified Negation Semantics?

    Negation does not carry a uniform interpretation across domains. In legal, regulatory, and medical reasoning, the intended interpretation depends on the reading in force -- open- versus closed-world, two- versus three-valued, and credulous versus skeptical. We study which reading…

  27. arXiv cs.LGSep 23Open code
    Geospatial embeddings detect old-growth forests but buffered spatial validation narrows their advantage over Sentinel features

    Old-growth forests develop over centuries under minimal anthropogenic disturbance, producing structurally complex and biodiverse stands. In Europe, protecting them requires mapping that is accurate for individual forest parcels yet deployable continent-wide. Geospatial foundation…

  28. arXiv econ.GNSep 23
    Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer

    Consumers increasingly delegate purchasing decisions to Large Language Models (LLMs) acting as surrogate consumers. Using "Tool-Lab," an adaptation of information-board process tracing that places product attributes behind costly tool calls, we examine how marketing pricing cues…

  29. arXiv cs.CVSep 23Open code
    Field-of-View Extension in Dental Cone-Beam CT via Implicit Neural Representations and Diffusion Model-Based Refinement

    Dental cone-beam computed tomography (CBCT) systems often employ detector configurations that provide a truncated field of view (FOV) that only captures a small part of the patient's anatomy. In this work, we aim to reconstruct an extended FOV using projections of truncated FOV s…

  30. arXiv cs.CLSep 23Open code
    Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models

    Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results. Reliable evaluation is particularly challenging for base models, whose limited i…

Columns

The people building AI, writing on their own blogs. Last 30 days.

  1. swyx (Latent Space)Sep 23
    🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)

    Radical Numerics is using biological chain-of-thought and multimodal perception to keep up with the bio-defense arms race, design new genomes and gain insights into biology itself.

  2. swyx (Latent Space)Sep 23
    [AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50%

    overshadowing more efficient GPT6 models from OpenAI

  3. Simon WillisonSep 22
    Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

    Yesterday was Grok 4.7 ( pelicans ) and MiMo v2.6 Flash/Pro ( more pelicans ). Today Anthropic released Claude Opus 5.5, and around an hour later OpenAI released GPT-6 Sol and GPT-6 Luna. It's going to take a while to get a good read on all of these new models, but here are my im…

  4. swyx (Latent Space)Sep 22
    🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science

    We talked to Google’s Oscar winning “Giganerd” about automating science, solving climate change, and how future generations can contribute to science in the age of superintelligent AI

  5. Sebastian RaschkaSep 22
    MiMo-V2.6 Pro Architecture and Training Notes

    Notes on MiMo-V2.6 Pro's GQA and sliding-window attention, agent training tasks, reward signals, and large RL batches.

  6. Nathan LambertSep 22
    Debating RSI, the US-China Gap, and Jaggedness with JS Denain of Epoch AI

    Podcast #19

  7. Simon WillisonSep 21
    Jev introduces a new shape of LLM - System One, aka Decision Models

    Last week TypeSafe AI unveiled Jev, their first example of a new category of model that they are calling "System One models" (I'm with Maggie Appleton, I think "decision models" is a better name for these). Jev is an interesting variant on the usual LLM format: it still accepts t…

  8. Nathan LambertSep 21
    The current balance of power in open models

    The expanded form of a testimony I prepared for Congress.

  9. Sebastian RaschkaSep 20
    It's Easy to Dismiss Jev as Just a Classifier

    A short note on Jev's generalization, possible encoder-style architecture and training, and Choice and Noul API examples.

  10. Thorsten BallSep 20
    Joy & Curiosity #100

    Interesting & joyful things from the previous week

  11. Nathan LambertSep 19
    Why I still haven’t bought into true RSI

    An “AI moderate’s” view on recent events and the trajectory of frontier models.

  12. Ethan MollickSep 18
    The Overhang

    Using your deep knowledge, wide knowledge, taste, and agency

  13. Hamel HusainSep 18
    AI Evals: Everything You Need to Know

    This document curates the most common questions Shreya and I received while teaching 5,000+ engineers and PMs AI Evals. Warning: These are sharp opinions about what works in most cases. They are not universal truths. Use your judgment. How to use this FAQ Browse the questions tha…

  14. Dan AbramovSep 17
    How I Vibed a Proof of Conway’s Conjecture

    You can just prove things, apparently.

  15. Dwarkesh PatelSep 17
    Noam Brown – Agent swarms, alignment, & recursive self-improvement

    “We never want to be in a situation again where we underestimate the AI.”

  16. Martin FowlerSep 17
    I don't like LLMs

    I have a lot of mixed feelings about AI and LLM technology. I’m fascinated by its effect on our profession, excited by the potential gains in productivity - and thus the products we could rapidly build. On the other hand, I’m fearful of the damage AI might cause: agent swarms tak…

  17. Martin FowlerSep 16
    Fragments: September 16

    Reports of agentic hacking continue, in this case it happened back in May and it seems OpenAI did not disclose that they were responsible. Simon Willison sees two options: After the Hugging Face and Wiki attacks OpenAI were still unable to review their previous logs and determine…

  18. Martin FowlerSep 15
    Nail the Narrative

    Sumeet Gayathri Moghe finds many folks building presentations get tangled in building slides without a coherent narrative. He advises distilling the big idea, visualizing the audience, and building a structured storyline. more…

  19. Sebastian RaschkaSep 14
    Pacing != Pacing Development

    My take on AI model pacing as a framework for release checks and the competitive pressure around model releases.

  20. Armin RonacherSep 13
    Interpreting Pangram

    Yesterday David Sacks wrote a tweet and within a few minutes people did, what they usually do, and they asked Pangram if it was AI. And Pangram said it’s entirely AI generated. To which David replied that these AI detectors are bogus. Now Pangram has a pretty low false posi…

  21. Addy OsmaniSep 13
    Brownfield Agentic Engineering

    What it takes to run agents in a codebase older than the team

  22. Simon WillisonSep 12
    Generating running routes with GPT-6 Astra and ChatGPT Work

    Here's a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at <my address>. Figure out 5K and 10K running routes from me that loop from my house. Use OSM data. It worked for 27 minutes and produced exactly what I'd asked for, as both an embedded v…

  23. Thorsten BallSep 12
    Joy & Curiosity #99

    Interesting & joyful things from the previous week

  24. Simon WillisonSep 11
    OpenAI agents attacked RubyGems back in May

    OpenAI agents carried out an undisclosed attack on RubyGems is a new bombshell report from Spencer Kitts, Thomas Larsen, and Sydney Von Arx - three of the four authors of the report on the agent attack on disused wikis ( previously ) last week. This time they're noting that it lo…

  25. Armin RonacherSep 11
    P(doom)

    This week some flavor of “AI is going to kill us all” went viral. In particular one where an employee put his personal probability of that happening above 10%. Which made me go to the Wikipedia page of P(doom) and I realized that Dario Amodei’s apparent probabil…

  26. Dwarkesh PatelSep 11
    AI researchers debate how close we are to recursive self-improvement

    “We're nowhere near the ceiling.”

  27. Simon WillisonSep 8
    Some thoughts on the Navier–Stokes Millennium Prize Problem

    On the Navier–Stokes Millennium Prize Problem introduces an impressive result from OpenAI, who used an unreleased model to produce a resolution to the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems that have been subject to a $1,000,000…

  28. Dwarkesh PatelSep 8
    Pretraining progress is mostly coming from data

    Breaking down 6 years of pretraining progress into data vs model improvements

  29. Armin RonacherSep 6
    Astra for Coding: Why Are We Doing This Again?

    I’m more and more convinced that all of AI engineering is Neijuan (内卷, meaning curl inwards). In China it describes a system that demands ever more effort and competition without improving output. The way in which it sometimes shows up in the West is the 996 nonsense. The E…

  30. Thorsten BallSep 6
    Joy & Curiosity #98

    Interesting & joyful things from the previous week

  31. Ethan MollickAug 30
    Agency and Agents

    From the Hugging Face Incident to Twilight Factories

  32. Addy OsmaniAug 30
    Agentic Skill Decay

    Agents can finish the task without teaching you anything. Building expertise now has to be deliberate.

  33. Addy OsmaniAug 26
    Audit your Agent files

    A practical guide to auditing what your coding agent still needs.

Technical Guides

Tutorials, case studies and technical deep dives.

  1. NVIDIA DeveloperSep 23
    Validate GPU Cluster Readiness Before AI Workloads Land
  2. Hugging FaceSep 23
    How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows
  3. AWS Machine LearningSep 23
    From portal-hopping to instant answers: HEMA’s journey with MCP and Amazon Bedrock
  4. AWS Machine LearningSep 23
    Agentic conversational video intelligence built on AWS
  5. AWS Machine LearningSep 23
    Use open weight models as your AI coding agent with Amazon Bedrock
  6. NVIDIA DeveloperSep 23
    How SWE-Serve Exposes the Gap Between Local Tests and Live Serving
  7. DatabricksSep 23
    How Concurrence governs clinical AI at a trillion-token scale with Unity Gateway
  8. OpenAISep 23
    Ringg’s AI agents resolve up to 65% of customer calls with OpenAI
  9. LangChainSep 22
    What Is Jev? A Guide to TypeSafe AI’s System One Model
  10. Apple Machine LearningSep 22
    How to Guide Your Language Flow
  11. Together AISep 22
    How to train your own Jev for $17
  12. VS CodeSep 22
    Building the new GitHub Copilot Inline Suggestions Model: Part Two
  13. ModalSep 22
    How to serve trillions of tokens for trillion-parameter coding agents
  14. DatadogSep 22
    Teaching a 9B model to investigate production alerts
  15. OpenAISep 22
    Better prompt caching for GPT-6
  16. NVIDIA DeveloperSep 22
    Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing
  17. AWS Machine LearningSep 22
    Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore
  18. NVIDIA DeveloperSep 22
    Topology-Aware Workload Scheduling with NVIDIA Topograph
  19. LangChainSep 22
    The Reliability Layer for Healthcare AI: Common LangSmith Use Cases
  20. AWS Machine LearningSep 22
    How Reactiv automates mobile commerce 80% faster with Amazon Bedrock AgentCore
  21. AWS Machine LearningSep 22
    Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI
  22. AWS Machine LearningSep 22
    How Trane gets building insights 60x faster with Amazon Bedrock AgentCore
  23. AWS Machine LearningSep 22
    How Tata Elxsi detects industrial safety risks in seconds on AWS
  24. AWS Machine LearningSep 22
    Extending public sector intelligence with Agentforce and AWS
  25. OpenAISep 22
    Parallel cut research time and cost in half with GPT‑6 Astra
  26. NVIDIA DeveloperSep 22
    Accelerating a ROS 2 Node with an AI Agent and NVIDIA Isaac ROS
  27. Hugging FaceSep 21
    How UK AISI and EvalEval Are Making Benchmark Results Reproducible
  28. Together AISep 21
    Canary rollouts: upgrade models in production without downtime
  29. WeaviateSep 21
    Agent Memory with Engram: A Practical Guide
  30. GitLabSep 21
    How GitLab reduced code-per-agentic-flow ratio by 45%
  31. DatadogSep 21
    Cut AI agent cost and improve accuracy with Code Execution in the Datadog MCP Server
  32. DatadogSep 21
    When users don’t click thumbs up: Inferring agent feedback from Datadog telemetry
  33. Shopify EngineeringSep 21
    Helix: The internal tool powering our Shopify app's native migration
  34. NVIDIA DeveloperSep 21
    How to Evaluate AI Agents From Tool Calls to Task Completion
  35. LambdaSep 21
    From structures to dynamics: scaling AI for molecular dynamics in drug discovery
  36. AWS Machine LearningSep 21
    Run Positron on Amazon SageMaker AI for data science workflows
  37. AWS Machine LearningSep 21
    How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore
  38. LangChainSep 21
    Can Jev Be a Better Agent Evaluator?
  39. NVIDIA DeveloperSep 21
    Turn Your Latest Observations Into Timely Weather Decisions With NVIDIA Earth-2
  40. Hugging FaceSep 21
    Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem
  41. SentrySep 21
    Building Sentry's Laravel AI Integration
  42. OpenAISep 20
    How V7 gives AI agents institutional memory
  43. Hugging FaceSep 20
    tokenizers v1: encode, decode and scaling, measured
  44. AWS Machine LearningSep 18
    Amazon SageMaker Inference: 2026 year-to-date launches in review
  45. NVIDIA DeveloperSep 18
    Benchmarking LLM Inference at Scale with AIPerf
  46. AWS Machine LearningSep 18
    Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime
  47. Together AISep 17
    How a global fintech scaled coding agent traffic with Dedicated Model Inference
  48. WarpSep 17
    Using LLM-as-a-judge scoring to measure your software factory
  49. LM StudioSep 17
    Splash Engine - the fastest local Qwen3.8 on Apple Silicon
  50. LangChainSep 17
    How Included Health Built Federated Healthcare Agents with LangGraph and Deep Agents
  51. DatabricksSep 17
    Database for AI Agents: 5 Evaluation Criteria
  52. Airbnb TechSep 17
    The guest journey, updated in real time: extending Airbnb’s sequence recommender with Chronon
  53. JetBrainsSep 17
    Building a RAG Pipeline for Semantic Code Search: A Developer Diary and Field Notes
  54. LM StudioSep 16
    Session References and Introspection
  55. Airbnb TechSep 15
    Beyond the model: Engineering AI infra with scientific judgement
  56. Pinterest EngineeringSep 11
    Evolving Pinterest’s Embedding Retrieval Platform
  57. Pinterest EngineeringSep 10
    Building Pinterest’s VLM Serving Stack on NVIDIA Dynamo
  58. CMU Machine LearningAug 10
    Forking-Sequences — Part II: Multi-Horizon Forecast Ensembling with Reduced Volatility
  59. Berkeley AI ResearchJul 29
    From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

Releases

Stable versions of the tools and libraries we track. Prereleases are left out.

ProjectVersionDate
Ollama v0.34.4 Sep 23
llama.cpp b11149 Sep 23
llama.cpp v0.5.0 Sep 23
llama.cpp b11147 Sep 23
LangChain langchain-anthropic==1.7.4 Sep 23
LangChain langchain-openai==1.6.5 Sep 23
LangChain langchain-openai==1.6.4 Sep 22
LangChain langchain-anthropic==1.7.3 Sep 22
Ollama v0.34.3 Sep 22
vLLM v0.30.0 Sep 22
LangChain langchain-fireworks==1.6.2 Sep 22
LangChain langchain-deepseek==1.1.1 Sep 22
LangChain langchain-openrouter==0.2.9 Sep 22
LangChain langchain-openai==1.6.3 Sep 21
LangChain langchain-core==1.6.4 Sep 21
Ollama v0.34.2 Sep 17
vLLM proto-v0.3.0 Sep 17
Open at the source ↗