OpenAI’s Decisions API enters public beta, returning typed answers about 10x faster at $0.10 per 1M input tokens.
Google Cloud’s Gemini agent is one enterprise agent routing work across Gemini and Claude models, with spend caps.
NVIDIA's PivotOPD trains multi-turn agents to recover from pivotal mistakes and beats 13 baselines on ALFWorld, WebShop, ...
Claude Haiku 5.5 brings 1M context, adjustable effort, and $0.10 input pricing to high-volume subagent and browser workloads.
JetBrains Mellum2.1 is an open 12B MoE coding agent model with 2.5B active parameters and 47.0% SWE-bench Verified.
Liquid AI has released Open d1, two open-weight multimodal models in its d1 decision model family. d1-3B reads text and ...
Unsloth details how Studio scans model code, blocks flagged weights, inspects packages and sandboxes tools before anything ...
Mistral AI released Mistral Large 4, a 1.05T parameter open-weight multimodal MoE with 49B active parameters, 1M context.
Reka Rho-1: a 19B omni-reasoning model unifying text, image, video generation and robot actions inside one shared network.
Architect's Liquid Inference auctions every LLM request across competing providers, locking a max price before the first ...
Reflection AI's Beam is a 501B open-weight MoE with 23B active parameters, 1M context, and Apache 2.0 weights.
Google DeepMind's EmbeddingGemma 2 embeds text, code, images, video and audio in one 740M model for on-device RAG.