Composer 2.5 is now available inside Grok Build.
Composer 2.5 is a fast, highly intelligent model that excels on long-running tasks and following complex instructions.
Introducing Qwen3.7-Plus — a multimodal agent model that unifies vision and language into one versatile agent foundation.
Multimodal interactive hybrid agent: unified GUI & CLI operation across visual and text tasks
Versatile coding agent & productivity assistant with full-modality input
Visual Agent: perception, reasoning, grounding, and search-augmented QA
Cross-harness generalization across diverse agent frameworks
One model. Sees, thinks, codes, acts.
Now available via API on Alibaba Cloud Model Studio. Try it — let us know what you build.
Blog:
https://
qwen.ai/blog?id=qwen3.
7-plus
…
Qwen Studio:
https://
chat.qwen.ai/?models=qwen3.
7-plus
…
API:
https://
modelstudio.console.alibabacloud.com/ap-southeast-1
?tab=doc#/doc/?type=model&url=2840914_2&modelId=qwen3.7-plus&serviceSite=international
…
Anthropic's latest update to Claude will allow the AI chatbot to generate custom charts, diagrams, and other visualizations during your conversation. If Claude determines a visual is useful based on the context of your chat, it will insert the image in-line, rather than in its side panel.
As an example, Anthropic says a conversation about the periodic table could lead Claude to generate a visualization of it, featuring interactive elements that let you click inside the table for more informat...
Anthropic rolled out two significant Claude updates this week. First, Claude can now build interactive charts and diagrams directly inside the chat window — available in beta on all plans including free. Second, Claude for Excel and Claude for PowerPoint now share full conversation context when multiple files are open, letting users pull data from spreadsheets into presentations without manually switching tabs. The dual release reinforces Anthropic's push into enterprise productivity workflows. Separately, the Ramp AI Index flagged Anthropic as the top AI stack choice for businesses, adding third-party validation to the product momentum.
Google launched Immersive Navigation, a new 3D navigation mode for Google Maps that Sundar Pichai called the product's biggest upgrade in over a decade. The view renders a vivid, real-time 3D picture of your surroundings with road-level details including lane markings and crosswalks. The update is part of a broader reimagining of Google Maps that the team described as built for the Gemini era, using AI to bring richer context and spatial understanding to everyday navigation. Logan Kilpatrick, who sat down with the Maps team, called it an impressive demonstration of Gemini in action at product scale.
Today we’re talking about the messy, fast-moving situation at Anthropic, the maker of Claude that now finds itself in a very ugly legal battle with the Pentagon.
The back-and-forth is complicated, but as of a few days ago, the Pentagon had deemed Anthropic a supply chain risk, and Anthropic has filed a lawsuit challenging that designation, saying the government has violated its First and Fifth Amendment rights by “seeking to destroy the economic value created by one of the world’s fastest-gr...
The impact of artificial intelligence extends far beyond the digital world and into our everyday lives, across the cars we drive, the appliances in our homes, and medical devices that keep people alive. More and more, product engineers are turning to AI to enhance, validate, and streamline the design of the items that furnish our…
Multiple AI agent companies announce major funding rounds, reflecting investor enthusiasm for autonomous AI systems. Notable raises include MultiOn ($35M), Adept ($150M), and Imbue ($200M) for agent development platforms.
The former Tesla AI director releases a comprehensive free course covering neural networks from scratch. Topics include backpropagation, transformers, and LLM training, with hands-on coding exercises in Python.
Zapier introduces AI-powered automation that understands natural language instructions. Users can describe workflows in plain English, and the AI builds and configures the appropriate Zaps across 6,000+ integrated apps.
SD4 brings photorealistic image generation and 4-second video clips from text prompts. The model shows significant improvement in text rendering and human anatomy, addressing long-standing issues with previous versions.
Windows 12 preview showcases Copilot deeply integrated into the OS, with ability to control settings, manage files, and automate workflows. New "Recall" feature provides photographic memory of user activity for instant retrieval.
Perplexity Enterprise allows companies to connect internal documents, databases, and wikis for AI-powered search. New features include role-based access control, audit logs, and integrations with Slack, Notion, and Google Workspace.
The AI-powered code editor Cursor secures major funding to expand its team and capabilities. The company reports 2M+ active developers and plans to introduce collaborative coding features and enterprise security controls.
SmolVLM enables multimodal AI on smartphones and IoT devices with models under 2B parameters. The release includes optimized versions for iOS and Android, bringing vision capabilities to mobile apps without cloud dependencies.
Meta's Llama 4 family includes models from 8B to 400B parameters, with the largest variant matching GPT-4 on most benchmarks. Released under permissive license for commercial use, marking a significant milestone for open source AI.
Gemini 2.5 Pro introduces native multimodal understanding across text, images, audio, and video. New agentic capabilities allow the model to perform complex tasks autonomously, including research, data analysis, and content creation.
Claude 3.7 Sonnet sets new records on SWE-bench, solving complex software engineering problems with 62% accuracy. The model introduces enhanced tool use capabilities and improved instruction following for enterprise workflows.
OpenAI announces GPT-4.5, featuring significant improvements in mathematical reasoning and code generation. The new model demonstrates 15% better performance on MATH benchmark and supports longer context windows up to 256k tokens.
Augment Code argued that modern development is shifting away from classic IDE assumptions and toward workspaces where developers define intent and delegate execution to agents. The company said the basic unit of interest is no longer a single file but an agent, suggesting the next generation of developer tooling will be organized around orchestration rather than manual code navigation. It is a notable framing because it pushes the coding-agent conversation beyond model quality into the shape of the actual interface developers may end up using every day.
Replit announced it raised $400 million at a $9 billion valuation, with investors including Georgian and G Squared. On the same day, it launched Replit Agent 4, featuring real-time multi-user collaboration — multiple people building in the same workspace simultaneously — and a new canvas mode that renders live app previews inline while you code. Early testers described the leap as the biggest product improvement they had felt in any tool. The timing of the funding and launch together signals Replit is positioning itself as the primary platform for AI-native software creation.
Lightning AI promoted Nvidia's Nemotron 3 Super as a model developers can customize, fine-tune, and deploy for reasoning agents in minutes, while related posts from Nvidia and Artificial Analysis emphasized the model's open weights, efficiency, and launch-day availability across inference providers. Taken together, the posts frame Nemotron 3 Super as more than another model release: it is being positioned as an open reasoning model with a real deployment ecosystem already wrapped around it. That combination of openness, benchmark credibility, and immediate infrastructure support is what gives the launch its weight.
Kaggle said its Community Benchmark SDK now supports automatic tracking for token usage, cost, and latency alongside standard evaluation results. The update pushes benchmark workflows closer to real product decisions, where teams need to understand not just which model performs best, but which one is cheapest and fastest to run. It is a useful signal that practical model economics are becoming part of mainstream benchmark tooling rather than an afterthought.
Google has completed its acquisition of Wiz, the cloud security company, with Sundar Pichai welcoming the Wiz team publicly. The deal gives Google a broad cloud security platform that protects workloads across providers, strengthening its pitch to enterprise customers who run multicloud environments. Pichai framed the acquisition as giving customers a comprehensive platform to secure their cloud and AI workloads. Wiz co-founder Assaf Rappaport previously turned down a $23 billion offer from Google, making this closing a notable reversal and one of the largest cybersecurity acquisitions in recent memory.
Rakuten uses Codex, the coding agent from OpenAI, to ship software faster and safer, reducing MTTR 50%, automating CI/CD reviews, and delivering full-stack builds in weeks.
Feng Qingyang had always hoped to launch his own company, but he never thought this would be how—or that the day would come this fast. Feng, a 27-year-old software engineer based in Beijing, started tinkering with OpenClaw, a popular new open-source AI tool that can take over a device and autonomously complete tasks for a…
Wayfair uses OpenAI models to improve ecommerce support and product catalog accuracy, automating ticket triage and enhancing millions of product attributes at scale.
How OpenAI built an agent runtime using the Responses API, shell tool, and hosted containers to run secure, scalable agents with files, tools, and state.
Abacus.AI CEO Bindu Reddy said her team is racing to switch coding workloads to GPT-5.4 because it performs better on fairly complex codebases and harder problems. In parallel, Augment Code said GPT-5.4 had become the default model in its agent development environment and framed it as especially strong for agent coordination. Taken together, the posts suggest GPT-5.4’s momentum is no longer just about benchmarks: it is turning into real adoption pressure inside coding and agent-engineering workflows.
A widely shared post from trq212 said Claude Code now supports `/btw`, a command for starting side-chain conversations while Claude continues working on the main task. Jason Liu reposted the feature to his audience, helping turn it into one of the most visible coding-agent workflow updates in this scrape batch. The change points toward a more interruptible, multitasking model for agent-assisted development rather than a single linear prompt-and-wait loop.
AgentMail announced a $6 million seed round led by General Catalyst, and Yohei Nakajima amplified the raise by arguing that email is becoming a core layer for AI agents, not just a communications channel. In his framing, email gives agents identity, authentication, notifications, and access to account creation flows such as self-service API signup and key retrieval. The combined posts cast AgentMail as infrastructure for practical autonomous workflows rather than another narrow inbox tool.
Two posts from the last 48 hours point to GPT-5.4 gaining real traction in coding-agent workflows. Augment Code said GPT-5.4 is now the default model in Intent and highlighted it as especially strong for agent coordination, while Abacus.AI CEO Bindu Reddy said her team is racing to move coding workloads onto GPT-5.4 because it performs better on fairly complex codebases and hard problems. The takeaway is that GPT-5.4 may be moving beyond benchmark headlines into day-to-day developer tooling decisions.
Andreessen Horowitz says new per-capita usage analysis across the 10 largest LLM products shows the United States ranking only 20th in AI adoption despite building many of the category’s biggest products. The firm framed the finding as evidence that consumer AI is becoming a more global market than top-line traffic charts imply. If that view is right, AI distribution and user behavior may become at least as strategically important as model leadership in the next phase of consumer competition.
AgentMail says it has raised a $6 million seed round led by General Catalyst, with investors including Y Combinator, Paul Graham, Dharmesh Shah, and Matt Shumer. The pitch is simple but strategic: every AI agent needs its own inbox. If that framing sticks, email could become one of the core infrastructure layers for autonomous software, giving agents a native way to receive updates, authenticate into workflows, and interact with the outside world without piggybacking on a human account.
Techmeme highlighted Axios reporting that Nielsen’s Gracenote has sued OpenAI for copyright infringement, alleging that OpenAI copied Gracenote’s data and the relational framework it uses to connect metadata. The case is notable because it extends the AI copyright fight beyond training on expressive works into the structured data layers that help classify and link content. If courts take these claims seriously, the legal risk for AI companies could widen from scraped media itself to the metadata systems that make media usable at scale.
A post flagged by The Rundown AI says OpenAI has added interactive visuals for more than 70 math and science concepts inside ChatGPT, including variable sliders, live graphs, and animated demonstrations. The update matters because visual explanations can make ChatGPT more useful for education and self-study than text responses alone, especially for subjects where intuition comes from seeing systems change in real time. It is another sign that OpenAI is expanding ChatGPT from a chatbot into a broader interactive product surface for learning and problem-solving.
Y Combinator CEO Garry Tan said builders should not sleep on Groq paired with Llama 4 Maverick, describing the combination as very useful for low-latency tasks. The post is notable because real-time responsiveness remains one of the hardest constraints in production AI systems, especially for assistants and agent workflows where delays directly affect usability. Tan’s endorsement suggests the conversation is shifting from pure benchmark leadership toward which model-and-inference stacks actually feel fast enough to use continuously.
LangChain announced `langgraph deploy`, a CLI flow that deploys an agent to LangSmith Deployment with a single command. The release is notable because it targets a familiar pain point in the agent stack: moving from experiments into something teams can run and monitor in production without stitching together custom deployment steps. In effect, LangChain is packaging deployment as a first-class part of the LangGraph workflow rather than an afterthought for platform engineers.
Google AI Developers said Gemini Embedding 2 is now available in preview through the Gemini API and Vertex AI, describing it as the company’s most capable and first fully multimodal embedding model built on the Gemini architecture. Jeff Dean separately amplified the launch, saying the model brings text, images, video, audio, and documents into the same embedding space. The update matters because embeddings sit underneath search, retrieval, and recommendation systems, and a stronger multimodal option gives developers a more practical foundation for building cross-format AI products.
A Hugging Face post shared by Georgi Gerganov introduced Storage Buckets, a new S3-like object storage option on the Hugging Face Hub and the first new repository type the platform has added in four years. Unlike the Hub’s standard versioned repos, Storage Buckets are mutable and non-versioned, with pricing positioned below Amazon S3. The release is significant because it shows Hugging Face expanding from model distribution into underlying storage infrastructure for AI teams building production systems.
Dimillian said he is joining OpenAI at the end of the month and will work on Codex as part of the developer experience team. He said he plans to bring what he learned from building Codex Monitor into the role, signaling that OpenAI is continuing to invest not just in coding models themselves, but in the tooling and workflows around how developers use them. The post matters because it points to deeper product focus on making Codex more usable inside real software teams, where monitoring, feedback loops, and developer experience often determine adoption.
Notion introduced Number Charts, a new dashboard element for displaying a single metric with customizable threshold-based colors. In the company’s post, the feature is positioned as a fast way to let one number tell the story, with yellow, green, and red states for quick status scanning. The launch broadens Notion’s reporting and dashboard toolkit, giving teams a simpler way to surface KPIs inside shared workspaces.
Andreessen Horowitz said the latest edition of its Top 100 Gen AI Consumer Apps ranking shows how quickly the consumer AI market is evolving beyond a narrow set of chat products. The firm argued that the newest leaders are increasingly global, multimodal, and embedded in everyday workflows, while Erik Torenberg separately amplified the release as evidence that the category deserves a refreshed lens on usage. The post matters because a16z’s ranking is one of the most widely watched snapshots of consumer AI adoption, and this edition points to a market where sustained engagement and product diversity are starting to matter as much as raw novelty.
Vinod Khosla said the real bar for robotics is autonomous performance in production environments, not polished lab demos, while highlighting Rhoda AI as a startup that impressed him with strong results from remarkably little robot training data. He emphasized the company’s use of internet-scale video pretraining to build a physical prior before deployment, suggesting a path to more general robotic capability without relying on massive amounts of expensive robot-specific data. The post matters because it captures a shift in how leading investors are judging physical AI: not by whether a robot can complete a staged demo, but by whether it can generalize reliably in the messy settings where commercial value is actually created.
Noam Brown said the core recipe behind frontier reasoning models looks surprisingly similar to AlphaGo. In his framing, the pattern is: imitate large volumes of human data, scale inference-time reasoning, and then apply reinforcement learning to move beyond imitation. The post stands out because it offers a concise mental model for how modern reasoning systems are evolving, linking today’s chain-of-thought and test-time compute strategies back to a landmark earlier system.
Hume said it is open-sourcing TADA, a text-audio dual alignment model designed to generate text and speech in one synchronized stream. The company said the architecture is meant to reduce token-level hallucinations while improving response speed, two of the main constraints that have limited real-time voice agents. The post matters because open-source voice models have often lagged behind closed systems on reliability and interaction quality, and Hume is positioning TADA as a practical step toward production-grade spoken AI.
Posts amplified by Databricks accounts point to OfficeQA Pro as a benchmark designed to test grounded reasoning on realistic enterprise workflows, including finding documents, extracting values, and performing analyses. The key claim is that frontier agents still score under 50 percent end-to-end. If that result holds up, it suggests the gap between flashy reasoning benchmarks and dependable workplace automation is still much wider than the AI hype cycle implies.
Google DeepMind said that a decade after AlphaGo, the techniques pioneered in that system are still compounding across the company’s research stack. In a new retrospective, the lab said those methods have already been used to prove mathematical statements and to assist scientists in making new discoveries. The broader significance is that DeepMind is presenting AlphaGo not as a historical trophy but as an early foundation for agentic systems that can reason through hard scientific problems.
Dan Shipper shared a custom Codex skill that connects to PostHog and a production database, then scans product data to identify bottlenecks and actionable growth insights. He described it as a “growth investigator” that works surprisingly well, pointing to a broader pattern in this batch where agentic tools are creeping from code generation into product analysis and marketing operations. The idea matters because it hints at a next phase for coding agents: not just helping teams build software, but helping them diagnose why the software is or is not growing.