Curated daily AI news

AI Daily

Read-worthy AI news filtered from 34 sources. No fluff; just substantial launches, research, policy, tooling, and market moves.

1 active subscriber · daily curated delivery

Latest curated scan

AI Daily

Curated, read-worthy AI news only — filtered from 34 sources.

Research · arXiv cs.AI · Sep 3 · score 30

Multimodal Language Models as Text-to-Image Model Evaluators

arXiv:2505.00759v3 Announce Type: replace-cross Abstract: The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate T2I progress. We present Multimodal Text-to-Image Eval (MT2IE), an evaluation framework

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Research · arXiv cs.AI · Sep 3 · score 28

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

arXiv:2605.24661v4 Announce Type: replace Abstract: Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness, providing limited insight into how models reason, how reliably they behave under contextual variation, or how efficiently they reach conclusions. This paper proposes a unified m

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Research · arXiv cs.AI · Sep 3 · score 27

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds

arXiv:2609.02302v1 Announce Type: new Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluation can support. We present two techniques that make simulated alignment evaluations harder to distinguish from real deployments. Ou

Why read: Governance signal: useful for risk, safety, security, or policy context.

Business · The Verge AI · Sep 2 · score 27

Researchers fear safety disaster ahead of OpenAI’s Astra release

OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing. As details about the model trickle out, researchers are warning it "may be the single worst development for AI security/safety to date." Shortly after […]

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Analysis · The Decoder · Sep 3 · score 24

GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era"

OpenAI has released GPT-6 Astra, its most capable model yet. President Greg Brockman says it marks the start of the "AGI era." Astra tops benchmarks in math, coding, and cybersecurity and is the first model OpenAI rates as "critical" under its safety framework. During testing, it independently found two previously unknown zero-day vulnerabilities. The articl

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Research · MarkTechPost · Sep 3 · score 21

Meta AI Released Muse Spark 1.3: An Agentic Coding Model That Uses ~20% Fewer Tool Calls and ~25% Fewer Tokens Than Muse Spark 1.2

Perplexity has shipped hybrid compute for its Mac app, splitting a single Perplexity Computer task between frontier models in the cloud and a compact model running on the user's machine. Tasks start in the cloud for search, planning and reasoning, then hand sensitive steps down to the Mac without restarting or losing context. An on-device privacy gate decide

Why read: Builder signal: practical implications for developers and AI operators.

Analysis · The Decoder · Sep 2 · score 21

Gemini 3.8 Flash is Google's third budget model in six weeks while frontier models remain MIA

Google's Gemini 3.8 Flash, the third Flash model in six weeks, matches Claude Opus 5 on some agentic coding benchmarks at lower cost. But its "working harder" reasoning burns about 30 percent more output tokens per task, making it pricier in practice than its predecessor despite identical token rates. The article Gemini 3.8 Flash is Google's third budget mod

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Labs · OpenAI Blog · Sep 2 · score 20

Safety overview: GPT-6 Astra

GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.

Why read: Governance signal: useful for risk, safety, security, or policy context.

Research · Reddit ML · Sep 3 · score 19

Mol-JEPA - Multimodal molecular foundation model [R]

<!-- SC_OFF --><div class="md"><p>Hi everyone,</p> <p>I just quickly wanted to share a paper I was working on for around a year now. I created this summary website with key results: <a href="https://flogrammer.github.io/moljepa/">https://flogrammer.github.io/moljepa/</a></p> <p>TL;DR: its a multimodal JEPA model for molecules.</p> <p>There will be more work

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Business · AI News · Sep 3 · score 19

NVIDIA to acquire Hugging Face for $12.93B

NVIDIA has agreed to acquire Hugging Face for $12.93 billion to scale the open-source model repository’s platform and infrastructure. The transaction targets platform growth and infrastructure investment, aiming to expand AI access for enterprise developers, software engineers, and research institutions globally. Built over the past decade by Clem Delangue

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Analysis · The Decoder · Sep 3 · score 19

Meta closes in on the top with Muse Spark 1.3, and undercuts rivals on price

Meta has released Muse Spark 1.3, its fourth model in the series in five months. According to Artificial Analysis, the model gains the most on agentic benchmarks but still trails Claude Fable 5.1 and other top models. The strongest argument is price. At $0.55 per task, Muse Spark undercuts every comparably scored rival. The article Meta closes in on the top

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Research · MarkTechPost · Sep 3 · score 19

Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon

Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one model on one chip family, it averages 1.23x MLX-LM's prefill throughput and 1.35x its decode throughput on a 40-core, 128 GB M5 Max. The post Perplexity Open Sources Lily: A Rust + Metal Inference Engine f

Why read: Builder signal: practical implications for developers and AI operators.

Developer · InfoQ AI ML Data Engineering · Sep 3 · score 19

Cohere’s Parse 5 Promises Efficient Multi-Modal Information Extraction from Complex Documents

<img src="https://www.infoq.com/styles/static/images/logo/logo_bigger.jpg"/><p>Cohere has launched Parse 5, a multimodal foundation model designed to extract structured data from complex enterprise documents. The 2.3-billion-parameter system converts visually rich PDFs into Markdown while providing bounding box coordinates for visual grounding. It has been e

Why read: Product signal: a notable model or platform change worth tracking.

Research · Reddit ML · Sep 2 · score 19

I scraped 5.94 billion TikTok videos and 3.23 billion profiles in 3 weeks. Uploaded full dataset to Hugging Face for free. Step by step tutorial and code below. [P]

<!-- SC_OFF --><div class="md"><p>Just uploaded the full 5.94 billion TikTok video dataset to Hugging Face. It’s fully open source:<br/> <a href="https://huggingface.co/datasets/kuben-developer/tiktok-videos-4b">https://huggingface.co/datasets/kuben-developer/tiktok-videos-4b</a></p> <p>This dataset was collected using a TikTok mobile app reverse-engineeri

Why read: Builder signal: practical implications for developers and AI operators.

Research · Reddit ML · Sep 1 · score 19

Latent Reasoning Landscape in 2026: Mapping BDH-CQ, HRM/TRM, Coconut [D]

<!-- SC_OFF --><div class="md"><p>After following various arXiv papers and researcher discussions on X/bluesky about latent reasoning and continual learning, one idea which resonates strongly is that path forward (towards AGI) may depend less on generating ever-longer chains of thought and more on finding architectures that can reason beyond the token stream

Why read: Research signal: likely to contain reusable findings, benchmarks, or technical detail.

Research · MarkTechPost · Sep 3 · score 18

OpenAI Releases GPT-6 Astra: A 1.05M-Context Computer-Use Model Gated Behind a ‘Critical’ Cyber Threshold

OpenAI released GPT-6 Astra on September 3, 2026, positioning it as a computer-use flagship rather than a chat model. It reports 72.6% on OSWorld V2-Offline, replaces Codex compaction with searchable notes, and ships a 1.05M-token context at $10/$50 per million tokens. It is also the first OpenAI model to cross the Critical cybersecurity threshold, which sha

Why read: Governance signal: useful for risk, safety, security, or policy context.

Business · TechCrunch AI · Sep 2 · score 18

OpenAI’s new reasoning technique alarms AI safety experts

OpenAI’s new Astra model will use “recurrent depth,” a technique that allows the model to operate outside of the sequential thinking that characterizes most reasoning models.

Why read: Governance signal: useful for risk, safety, security, or policy context.

You are receiving this because you subscribed at http://ai.totaljerk.net. Unsubscribe link is included in subscriber emails.