Curated daily AI news

AI Daily

Read-worthy AI news filtered from 34 sources. No fluff; just substantial launches, research, policy, tooling, and market moves.

1 active subscriber · daily curated delivery

Latest curated scan

AI Daily

Curated, read-worthy AI news only — filtered from 34 sources.

Research · MarkTechPost · Sep 6 · score 27

UC Berkeley Researchers Release CUA-Lite, an Open Platform Unifying Sandboxes, Data, Evaluation and RL for Computer-Use Agents

Training and benchmarking a computer-use agent needs four things — agents, environments, traces, and a framework to evaluate and train them — and all four ship in incompatible formats today. CUA-Lite, from a UC Berkeley led team, puts them behind one action space and one data schema, and replaces OSWorld's per-task virtual machine with a plain Docker con

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Developer · InfoQ AI ML Data Engineering · Sep 5 · score 25

Beyond Zero: Google Publishes Successor to BeyondCorp

<img src="https://res.infoq.com/news/2026/09/google-beyond-zero/en/headerimage/generatedHeaderImage-1787655112500.jpg"/><p>In a recent research paper, Google introduced Beyond Zero, a “security model for the AI era” that extends Zero Trust to autonomous AI agents. The new approach moves access decisions from the application level to individual resources

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Research · MarkTechPost · Sep 7 · score 24

OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device

OpenBMB has released MiniCPM5-2B, a dense causal language model with 2,516,756,480 parameters and a native 131,072 token context. It averages 53.9 across the 34 benchmarks in its model card, ahead of Qwen3.5-4B at 51.1, with its clearest leads in tool use, coding agents and long-context retrieval. Post-training pairs 400B tokens of deep-thinking SFT with RL

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Research · MarkTechPost · Sep 7 · score 23

Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories

Robot datasets have grown far slower than the models trained on them, mostly because collection stays locked to lab hardware. AXIS moves demonstration collection into a web browser and pushes everything expensive to backend GPUs. The result is 207 tasks and 50,129 verified Franka trajectories, and continual pretraining that lifts π0.5 from 83.9 to 88.8 on L

Why read: Product signal: a notable model or platform change worth tracking.

Research · Reddit ML · Sep 5 · score 21

GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N]

<!-- SC_OFF --><div class="md"><p>A researcher has <a href="https://www.linkedin.com/posts/s-berezin_llm-aialignment-aisecurity-activity-7502013488412680192-IO6c/">reported</a> a jailbreak of GPT-6 Astra within a day after release.</p> <p>The attack is described as combination of TIP (Task-in-Prompt) attack from <a href="https://aclanthology.org/2025.acl-lon

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Analysis · The Decoder · Sep 8 · score 20

OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper

Mathematician Tristan Buckmaster says an OpenAI researcher pressured him after information about his AI-assisted progress on the Navier-Stokes equations allegedly reached the company. The researcher tried to remove his co-author because he works at Anthropic and threatened Buckmaster when he refused, according to Buckmaster's account. OpenAI then claimed its

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Infrastructure · AWS Machine Learning Blog · Sep 8 · score 20

Govern models with MLflow and Amazon SageMaker AI Model Registry sync: Part 1

Managed MLflow on Amazon SageMaker AI now syncs richer model metadata (training metrics, evaluation results, inference specs, and lineage) into the SageMaker AI Model Registry, with lifecycle stage promotion. Part 1 shows how to govern candidate models in a single account using IAM guardrails.

Why read: Builder signal: practical implications for developers and AI operators.

Developer · InfoQ AI ML Data Engineering · Sep 8 · score 20

GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access

<img src="https://res.infoq.com/news/2026/09/gitlab-ai-sandbox-access/en/headerimage/generatedHeaderImage-1788446396087.jpg"/><p>GitLab warns that isolating an AI coding agent in a sandbox does not necessarily make the agent safe. In a new security analysis, the company describes an internal evaluation in which an AI agent escaped its sandbox by exploiting a

Why read: Governance signal: useful for risk, safety, security, or policy context.

Infrastructure · AWS Machine Learning Blog · Sep 8 · score 18

Take on your most ambitious work with GPT-6 Astra on Amazon Bedrock

GPT-6 Astra from OpenAI is now generally available on Amazon Bedrock. It brings deeper reasoning and sharper judgment to your most demanding tasks, running on the Amazon Bedrock inference engine built for high performance, security, and scale.

Why read: Governance signal: useful for risk, safety, security, or policy context.

Labs · OpenAI Blog · Sep 8 · score 18

Funding grants for new research into AI and teen development

Apply now for OpenAI’s $5 million grant program supporting independent research on how generative AI affects teen development, well-being, and safety.

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Infrastructure · AWS Machine Learning Blog · Sep 8 · score 17

Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6

Benchmark two 30B Mixture-of-Experts models, Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B, across G5, G6, G6e, and G7 GPU instances on Amazon SageMaker AI. Compare throughput, latency, and cost-per-token, and see how G7's NVIDIA Blackwell GPUs deliver measurable price-performance gains for real-time LLM inference.

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Research · Reddit ML · Sep 5 · score 16

Astra vs. Fable 5.1 on real ML tasks -- tradeoffs, strengths, shortcomings [P]

<!-- SC_OFF --><div class="md"><p>I ran a side-by-side ML text-processing and model-training workflow using Fable 5.1 vs. Astra (both on xhigh), and the results could not have been more different. Warning, long post.</p> <p><strong>TL;DR -- Astra codes more agentically, Fable more coherently. Fable writes better and follows directions better. Astra&#39;s fin

Why read: Builder signal: practical implications for developers and AI operators.

Developer · KDNuggets · Sep 8 · score 15

5 Ways I Access Coding Models for Free

Explore five free ways to access AI coding agents, proprietary coding models, and open-weight models without paying for expensive subscriptions or GPUs.

Why read: Builder signal: practical implications for developers and AI operators.

Business · HackerNoon AI · Sep 8 · score 14

Why Can’t My AI Agent Renew My Driver's License?

Autonomous AI agents are transforming software engineering, but they crash into a wall at city hall. Between 40-year-old COBOL mainframes, prompt injection risks, and zero-tolerance administrative law, here is why government paperwork is the ultimate test for agentic tech.Read All

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Analysis · The Decoder · Sep 7 · score 14

How AI wiped out an entire industry in Nairobi

In Kenya, ChatGPT wiped out an entire business model: writing academic papers for foreign students. The article How AI wiped out an entire industry in Nairobi appeared first on The Decoder.

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Analysis · Import AI · Sep 7 · score 14

Import AI 472: DeepMind’s cheating math agents; populist AI policies; and Forethought theorizes a nightwatchman

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Researchers discover another OpenAI agent emergent communication incident:…Less severe, but worrying nonetheless…Some researchers recently found another incident of AI agen

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

You are receiving this because you subscribed at http://ai.totaljerk.net. Unsubscribe link is included in subscriber emails.