A newsletter that makes FOMO obsolete.
if i fear missing out on reading this newsletter, did i win or lose
01 · The One Thing
The 90-minute war
Tuesday morning, Anthropic launched Claude Opus 5.5: Fable-5.1-level performance, 30% faster output, framed as beating GPT-5.6 Sol on CursorBench at one-third the cost. Ninety minutes later, OpenAI dropped GPT-6 Sol and Luna and cut API prices 50% on the spot.
This is the first model cycle where the launch-day discourse is about unit economics first and capability second. The timeline is wall-to-wall price-per-task memes. Polymarket is literally pricing “up to 93% cheaper coding costs.”
Signal: If you build on APIs, your cost basis got rewritten twice before lunch on Tuesday. Every agent product’s margins just moved. Recalculate build-vs-buy before you ship anything priced on old token math. Sources: techcrunch on OpenAI · macrumors on Anthropic
02 · Lab Watch
GPT-6 Sol + Luna, prices halved
Sol for complex coding and agentic work, Luna for high-volume tasks. Sol now $2/$10 per 1M input/output tokens (was $4/$20), Luna $0.10/$0.50, cached reads 90% off. Permanent cuts, not promo.
Signal: OpenAI is passing inference-efficiency gains straight to developers. The Jevons playbook, live: cheaper thinking, more thinking.
Opus 5.5, and a November IPO looms
List prices $4/$20 per 1M tokens (20% below Opus 5), but cache reads dropped 60% to $0.20/1M. Anthropic’s pitch: cache dominates agentic workloads, so typical costs fall ~40%. Sonnet 5.5 and Haiku 5.5 due in weeks. First release since Amodei’s “pace the frontier” call.
Signal: Both labs cut prices the same day via different mechanisms. The real benchmark is now cost per completed agent task, not cost per token.
Grok 4.7 ships, Copilot day one
New largest base model, longer RL on multi-hour tasks, new safeguard stack, $2/$6 per 1M tokens. Rolling out to GitHub Copilot immediately. Independent Artificial Analysis testing matched xAI’s claims except Terminal-Bench 4.0: claimed 38.0%, measured 25.76%.
Signal: Vendor benchmarks remain guilty until proven innocent. Also: Copilot distribution gives xAI instant developer reach, which matters more than the benchmark gap.
Hardware blitz at Connect, plus a curtain pulled back
Zuckerberg unveiled Muse Charm (keychain-sized pocket device for the Muse agent, December), Ray-Ban Meta Gen 3 at $449, camera-free audio glasses at $349, an FDA-cleared hearing-aid feature inside glasses, and Muse computer-use for Mac. Meanwhile Reuters found human contractors secretly making Muse’s phone calls on users’ behalf, one of whom made a racist reference on a call. Meta apologized and rolled the feature back. Separately: Muse is already straining Meta’s compute at ~700K DAU, and Amazon blocked Muse from shopping on its platform.
Signal: The “fully autonomous agent” still has humans behind the curtain, at Meta scale. And agents eating 10x projected tokens is a brutal unit-economics problem nobody has solved. Sources: techcrunch on Connect · gizmodo on contractors
Gemini 3.8 Flash TTS takes voice benchmarks
New text-to-speech models: 100+ languages, 2,000 production voices, 30-second voice cloning, SynthID watermarking. Took #1 on Hume AI’s Voice Design Benchmark and Voice Arena preference evals.
Signal: Voice quality is now a leaderboard battleground, not a solved commodity. If your product has a voice, the frontier moved this week.
03 · Open Weights
NVIDIA Nemotron 3 Diarization
100M-parameter speaker diarization (who-spoke-when, up to 8 speakers), streaming and offline, reportedly #1 on VoiceArena’s diarization board.
Signal: Tiny, practical, open. This quietly becomes infrastructure in every meeting-recording pipeline. See “Build on This” below.
Black Forest Labs FLUX 3 Action
7B robotics world-action model predicting the next 32 robot actions from camera frames plus text. Claims 42.92% on NVIDIA’s RoboLab-120, beating NVIDIA’s own 16B Cosmos policy. FP8 variant runs on 24GB consumer GPUs. Commercial license terms unclear in coverage found.
Signal: Competitive robot policies on consumer GPUs could do for robotics what Stable Diffusion did for images. Verify the commercial license before building on it.
Alibaba Qwen-Image-2.1
7B open-weights text-to-image plus image editing, native 2K, day-zero support in Diffusers, ComfyUI, vLLM, SGLang.
Signal: Alibaba keeps shipping the most usable open image models. Day-zero ecosystem support means instant adoption.
Frontier watch
Moonshot’s Kimi K3 remains the first open model to top a board outright (#1, LMArena Frontend Code Arena). On Hugging Face, Qwen3.8-27B is now the most-liked model, with DeepSeek-V4-Pro and GLM-5.2 crowding the frontier band.
Signal: Open models now sit inside the frontier Elo band. The question shifted from “which model is best” to “which is best per dollar.”
04 · Money Flow
Island $400M at $6.4B · Tekever $580M at $6.4B
Same $6.4B valuation, same week, two theses: Island sells the enterprise browser as “scale AI agents securely” (Sequoia, Coatue, Insight in the round), Tekever sells AI drones weeks after a UK MoD surveillance contract worth up to £400M.
Signal: Money is flowing to AI with a government or enterprise-security buyer attached. Agent-security infra is commanding decacorn-adjacent valuations.
Ema raises $77M: “your SaaS apps become databases”
“AI employees” startup automating HR/IT/finance, led by Creaegis with Accel and Prosus. Valuation up 4x since 2024, $140M total raised.
Signal: The clearest articulation yet of the agent-eats-SaaS displacement thesis, now with revenue behind it. If you sell SaaS, this is your competitor’s pitch deck.
Go.AI raises $85M, profitable, on-prem
On-prem AI infra for regulated industries: 200+ customers, 8x ARR growth, 12.5M queries/day, profitable. Led by Updata Partners.
Signal: Sovereign AI is a real, profitable category. Banks, hospitals, and defense want frontier AI that never leaves the building.
Flatkey raises $10M: arbitraging the labs
Single API key across 100+ models and 1,000+ tools at 60-90% of official list prices. 10,000 developers in two months since July.
Signal: A startup growing this fast by undercutting list prices is itself a pricing signal. Margin compression is coming for undifferentiated model access.
05 · Leaderboard
Who’s king today
Claude Fable 5.1 holds the LMArena text crown at 1,514 Elo; Opus 5.5 debuted top-three on the Artificial Analysis Intelligence Index. Kimi K3 still the only open model to top a board outright. Seven of 86 vals.ai-evaluated models now clear 95% on SWE-Bench.
Signal: Benchmarks are saturating at the top, which is exactly why the labs pivoted the fight to price. Capability is table stakes; cost per task is the differentiator.
06 · The Argument
“Which one y’all taking”
AI Instagram and Threads are wall-to-wall with side-by-side price/performance charts from the 90-minute launch war: dueling AutomationBench and CursorBench claims, “up to 93% cheaper coding” memes, developers recalculating build-vs-buy in real time.
Signal: First cycle where economics lead capability in the discourse. When the crowd prices your infra like a commodity, build the thing on top, not the infra.
Altman and Amodei briefed the UN Security Council
Both CEOs jointly briefed the UN Security Council on AI risk on Sept 23, while Trump dismissed regulation talk at the General Assembly.
Signal: Mostly theater, but notable: the doomer-vs-accelerationist fight now has a diplomatic chapter. File under context, not action.
07 · Papers
Shopping by algorithm: how agentic AI deploys human heuristics as a surrogate consumer
Agents are becoming the buyers, and they shop with human-like heuristics.
Signal: If your checkout isn’t machine-readable, you’re invisible to the fastest-growing customer segment. The agentic-commerce wave is now a research topic, which means it’s about 12 months from being your problem.
Can LLMs reason about runtime behavior? A repository-level dynamic benchmark
Static benchmarks are saturating, so evaluation is moving to dynamic execution: can the model predict what code actually does when run?
Signal: Execution-based eval is the next frontier for coding agents. If you eval models, start tracking this line of work.
Agent-editing world model: rethinking world modeling for LLM agents
World models, but for LLM agents instead of robots: learning editable internal models of the environments agents act in.
Signal: Agents that simulate before acting beat agents that act and apologize. Worth a skim if you build multi-step agents.
Memory Attention
New attention variant aimed at long-horizon memory, landing the same week GitHub trending filled up with agent-memory repos.
Signal: Theory and practice converging on the same bottleneck: agents that can’t remember can’t compound. Read alongside hindsight and ai-memory below.
When and where to trust the teacher: unifying on-policy distillation and GRPO
Distillation theory catching up with practice: when teacher signals help, when they hurt, and how entropy mediates it.
Signal: Everyone distills now; few understand when it backfires. Practical if you train or fine-tune anything.
08 · GitHub Trending
The theme: memory, skills, and harnesses
Nearly every fast mover is agent infrastructure. The gold rush isn’t models anymore; it’s the scaffolding around them.
- vectorize-io/hindsight · +1,607 today · agent memory that learns
- google/ax · +1,376 today · Google’s open agentic orchestration runtime
- dream-num/univer · +1,060 today · office harness for agents
- alibaba/open-code-review · +9,833 this week · biggest mover
- affaan-m/ECC · +6,695 this week · harness perf optimization
- stablyai/orca · +6,435 this week · fleet-of-agents IDE
- Tencent/WeKnora · +4,522 this week · docs to queryable knowledge
- addyosmani/agent-skills · +3,867 this week · production agent skills
- JustVugg/colibri · +2,739 this week · frontier MoE in pure C
- akitaonrails/ai-memory · +1,208 this week · long-term memory in Rust
- obra/superpowers · +606 today · agentic skills framework
- superdesigndev/treg · +470 today · “OpenRouter for agent tools”
- strands-agents/harness-sdk · +463 today · build a harness end-to-end
- HKUDS/CLI-Anything · +415 today · make all software agent-native
Signal: Three bets the crowd is making: agents need memory (hindsight, ai-memory), agents need skills (superpowers, agent-skills), agents need cheaper/faster harnesses (ECC, ax). Google open-sourcing orchestration plumbing is the tell that this layer is commoditizing fast.
09 · Hacker News
What is RLCD? The secret behind Jev
The training method behind TypeSafe’s Jev, explained. 51 points, early thread.
Signal: The technical companion to this week’s Jev discourse. Worth the read.
ArXiv receives multiyear commitments as an independent nonprofit
The commons gets funded. ArXiv secures its future as independent infrastructure. 286 points.
Signal: Quietly important: the paper pipeline you rely on just got durable funding.
Hackers influence ChatGPT and Gemini to direct users to scam centers
Prompt injection at population scale: poisoned content steering assistants toward scam operations. 114 points.
Signal: The attack surface of “agents that browse” is now a crime beat. If you ship a browsing agent, this is your threat model.
Also notable
Agents.md speaks Unix, and you should too: agent conventions as unix philosophy, small file, big idea. The Year of Internal Tools: the quiet enterprise AI boom is internal tooling, not products. Oracle invokes force majeure on New Mexico AI data center: the buildout is hitting physical limits. LinkedIn wins court order blocking mass scraping: training-data moats just got legal precedent.
10 · OpenRouter Apps
Where the tokens actually flow
Ranked by tokens routed through OpenRouter. Coding agents own this chart: Hermes Agent (Nous Research, ~49T), Kilo Code (~13T), Cline (~12T), Freebuff (~7T). The top of the market is agentic coding, and it’s not close.
Signal: Whatever the labs announce, the revealed preference of developers is coding agents. Build distribution where the tokens are.
New and climbing
DecodingTrust Agent Platform (jevagent) appeared Sept 22 and is already ranked: third-party eval infra for Jev-style models, live within days of launch. Also new: sel-jev-rubric-judge (Sept 20), rustllm (Sept 18), LuckyRobots worker fleet (Sept 18).
Signal: Jev got an eval-harness ecosystem within 72 hours of launch. The Jevons paradox in real time: collapse the cost of thinking and new apps sprout before the paint dries.
11 · Watch
Yes, Jev Is Insane, But There’s A Catch
Signal: The clearest 10-minute technical take on Jev’s trade-offs.
Claude Opus 5.5 AI: An Incredible Leap Forward
An ex-OpenAI researcher just deleted language from the LLM…
Signal: That’s the TypeSafe story in Fireship’s 5-minute format: giving up string generation entirely. Good shareable explainer.
AI labs may be hiding their biggest breakthroughs
Elon Musk Just Released Grok 4.7 — And It’s NOT What We Expected
12 · Listen
Jev: System One models for Prod, not God — with Diogo Almeida
Latent Space · Sep 21 · episode
Signal: The founder interview behind the model.
Bio-security is an AI Arms Race
Latent Space · Sep 23 · episode
90 minutes of unfiltered product advice from Snap and Discord’s product leaders
Lenny’s Podcast · Sep 20 · episode
13 · Read
Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?
Lenny’s Newsletter · Sep 22 · read
Signal: The launch war, adjudicated blind. The most useful single comparison of the two models.
I left Claude for months. Opus 5.5 is why I’m back
Lenny’s Newsletter · Sep 22 · read
TLDR AI — Sep 24 edition: Gemini TTS, Claude’s novel enzyme, Google private memory
TLDR AI · Sep 24 · read
14 · Safe to Ignore
Permission to skip
Skip: The UN Security Council AI briefing. Diplomacy is not a product signal. Nothing announced, nothing decided. Skip: Bessemer’s $5.75B in new funds. Dry-powder press releases don’t change your roadmap. Note the direction of travel, then move on. Skip: Grok 4.7’s Terminal-Bench gap (38% claimed, 25.8% measured). Unless you benchmark models for a living, file under “vendor charts lie a little” and move on. Skip: F-Droid 2.0 and the Nokia Design Archive. Lovely threads. Not FOMO. Read them on the weekend.
15 · Build on This
The meeting-intel refinery
NVIDIA just open-sourced a 100M-parameter diarization model that does who-spoke-when and runs anywhere. Luna-class models now cost $0.10 per million input tokens. The two hardest parts of “upload a meeting, get intelligence out” just went to ~zero.
Product: drop in a recording, get chapters, action items, and a searchable transcript. Sell it as a $29 one-time credit pack for 10 hours of audio. The wedge: every incumbent charges per seat per month; this charges per meeting, once.
16 · Appendix
Everything below didn’t make the cut. Tell me what I promoted wrongly or buried unfairly and issue #2 gets sharper.
Hacker Newstop 30
- F-Droid 2.0
- Show HN: Whiteboard (YC W26) – An open-source IDE for thoughtful software design
- Two-tier encryption in the UK
- Agents.md speaks Unix, and you should too
- Nokia Design Archive (2025)
- WaveDigger: Dig into wireless signals to discover their physical locations
- GitHub has not removed malicious imitation software after 3 weeks
- B5-BJ2 – Ice Cream Barges – Concrete Ship Constructors (2023)
- Linux support is coming to Snapdragon X2 series
- Experiencing writing at our recent Chinese calligraphy workshop
- The science of Monkey Island: can grog dissolve a metal mug that fast?
- Ideas on modernizing the open-source desktop
- Federal judge orders Texas to air condition all prisons by the end of 2029
- RAM: the forgotten history (2024)
- When the Debugger Lies
- ArXiv receives multiyear commitments to support it as an independent nonprofit
- Enjoy Every Sandwich
- Coulomb's law remains tricky to test at home
- The newest ESP32 can run Linux and it's getting close to a Raspberry Pi
- VSCode's SSH Agent Is Bananas (2025)
- What Is RLCD? The Secret Behind Jev
- Contrastive Language Models
- The "Windows XP Box" (2003)
- Oracle invokes force majeure on New Mexico AI data center
- The Year of Internal Tools
- Fixing the Portobello Police Station Clock
- Hackers influence ChatGPT and Gemini to direct users to scam centers
- LinkedIn wins court order blocking mass scraping of user data
- Why 'What's Opera, Doc?' looks like that
- Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
GitHub trending · daily14 repos
- rohitg00/ai-engineering-from-scratchLearn it. Build it. Ship it for others.
- vectorize-io/hindsightHindsight: Agent Memory That Learns
- dream-num/univerThe Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.
- google/axGoogle's open agentic orchestration runtime
- NVIDIA/Model-OptimizerA unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
- FxEmbed/FxEmbedFix X/Twitter and Bluesky embeds! Use multiple images, videos, polls, translations and more on Discord, Telegram and others
- anthropics/financial-services
- HKUDS/CLI-Anything"CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
- mvt-project/mvtMVT (Mobile Verification Toolkit) helps with conducting forensics of mobile devices in order to find signs of a potential compromise.
- obra/superpowersAn agentic skills framework & software development methodology that works.
- strands-agents/harness-sdkBuild an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
- julyx10/lapAn offline-first photo manager for large local libraries
- superdesigndev/tregOpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn
- leejet/stable-diffusion.cppDiffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
GitHub trending · weekly20 repos
- anthropics/financial-services
- alibaba/open-code-reviewSecure, fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.
- anthropics/claude-codeClaude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
- affaan-m/ECCThe agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
- Tencent/WeKnoraOpen-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
- addyosmani/agent-skillsProduction-grade engineering skills for AI coding agents.
- anthropics/knowledge-work-pluginsOpen source repository of plugins primarily intended for knowledge workers to use in Claude Cowork
- stablyai/orcaOrca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
- TencentCloud/OctopA smarter, self-hosted AI assistant — multi-user, multi-agent.
- davila7/claude-code-templatesCLI tool for configuring and monitoring Claude Code
- JustVugg/colibriRun frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
- vectorize-io/hindsightHindsight: Agent Memory That Learns
- cloudflare/quiche🥧 Savoury implementation of the QUIC transport protocol and HTTP/3
- Fission-AI/OpenSpecSpec-driven development (SDD) for AI coding assistants.
- akitaonrails/ai-memorySolution for long term memory for agent coding CLIs and to facilitate handoff between different agent vendors
- pytorch/pytorchTensors and Dynamic neural networks in Python with strong GPU acceleration
- cline/clineAutonomous coding agent as an SDK, IDE extension, or CLI assistant.
- LibreChat-AI/LibreChatEnhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API, Azure, Groq, o1, GPT-5, Mistral, OpenRouter, Vertex AI, Gemini, Artifacts, AI model switching, message search, Code Interpreter, langchain, DALL-E-3, OpenAPI Actions, Functions, Secure Multi-User Auth, Presets, open-source for self-hosting. Active
- superdesigndev/tregOpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn
- cilium/ciliumeBPF-based Networking, Security, and Observability
Hugging Facemost liked · 30
- Qwen/Qwen3.8-27B
- black-forest-labs/FLUX.1-dev
- deepseek-ai/DeepSeek-R1
- moonshotai/Kimi-K3
- stabilityai/stable-diffusion-xl-base-1.0
- meta-llama/Llama-3.1-8B-Instruct
- CompVis/stable-diffusion-v1-4
- hexgrad/Kokoro-82M
- meta-llama/Meta-Llama-3-8B
- openai/whisper-large-v3
- sentence-transformers/all-MiniLM-L6-v2
- black-forest-labs/FLUX.1-schnell
- Qwen/Qwen3.8-Flash-Next
- MiniMaxAI/MiniMax-H3
- deepseek-ai/DeepSeek-V4-Pro
- Tongyi-MAI/Z-Image-Turbo
- openai/gpt-oss-120b
- zai-org/GLM-5.2
- meta-llama/Meta-Llama-3-8B-Instruct
- openai/gpt-oss-20b
- bigscience/bloom
- stabilityai/stable-diffusion-3-medium
- Lightricks/LTX-2.5
- meta-llama/Llama-2-7b-chat-hf
- mistralai/Mixtral-8x7B-Instruct-v0.1
- unsloth/Qwen3.8-27B-GGUF
- meta-llama/Llama-2-7b
- baidu/Unlimited-OCR
- deepseek-ai/DeepSeek-V3
- mistralai/Mistral-7B-v0.1
arXivcs.AI / cs.CL / cs.LG · newest 25
- On the Diffusibility of High-Dimensional LatentsRepresentation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reco
- Contrastive Learning for Authorship VerificationOur results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context length, and random text span data augmentation as important factors
- StudentBench: AI and human tutoring yield equivalent GRE learning gainsArtificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection
- Where Should I Join? Robot Group Joining via Language-Guided Goal PredictionSocial navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity and formation. This is a highly semantic task, yet an important capability for ap
- Even Sharper Bounds for Transductive Learning and Its ApplicationsWe introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--train empirical process. Its proof uses
- Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic BenchmarkLarge language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mos
- Nonequilibrium Phases of Repulsive Self-Attention: Chaos, Attention Condensation, and Emergent LocalityWe study the nonequilibrium dynamics of a minimal recurrent transformer with $N$ normalized tokens, $Q=K=I$, and a negative value map $V=-I$. Similarity-based attention selects nearby representations, while the negative value map drives tokens away from the selected field. This feedback can continua
- Order-Invariant Answers, Order-Sensitive Representations in Mathematical ReasoningReordering a set of mathematical rules without changing its meaning should preserve the correct answer, but must a model's internal representations stay invariant too? We investigate this question using synthetic multi-step function-composition problems, each presented under multiple rule orderings
- Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip ConnectionsWe study minimal-norm interpolation and $\ell_2$-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-lay
- Cross-Scale Transfer Learning for Depression Severity Prediction: From PHQ-8 to HAMD-17 Across Languages and Clinical ParadigmsThis work addresses continuous depression-severity score prediction from clinical interview transcripts under data scarcity. We propose a sequential low-rank adaptation (LoRA) protocol for cross-scale transfer: a Qwen3 backbone with a bounded regression head is first fine-tuned on the English DAIC-W
- Context-Continuous Preference Learning for Exoskeleton PersonalizationPersonalizing exoskeleton assistance across operating conditions is constrained by the time and physical effort required to collect user feedback. We examined whether a user's preference landscape varies smoothly across operating conditions and when this continuity supports learning from limited fee
- Repairability of Inexact Solvers in Recursive State Estimation with Machine LearningRecursive state estimation often executes approximate numerical solutions inside a feedback loop, where highly accurate local steps do not guarantee better overall results. For a fixed linear Kalman model, we characterize when a correction within a prescribed subspace and norm budget can meet a loca
- Agent-Editing World Model: Rethinking World Modeling for LLM AgentsRecent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool res
- Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World ModelLatent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow never moves the manipulated object; retraining it with latent-only losses only trades stillness for te
- Learning Holographic Reduced Representations with Clifford Variational AutoencodersVector Symbolic Algebras project data structures into a hyperdimensional vector space through the application of their vector algebras to randomly generated atomic vector symbols and fractional power encodings of real-valued data. Embedding unstructured data remains an open question. We present \tex
- Learning Collective Dynamics with Differentiable Gaussian RepresentationsCollective responses depend on individual differences, contact opportunities, and accumulated experience. Learning their dynamics from aggregate counts requires connecting a population's response distribution to both current observations and future behavior. We introduce Differentiable Gaussian Dyna
- Memory AttentionLanguage models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory A
- Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction FollowingFine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings transfer to machine translation (MT) fine-tuning and to MT-specific inst
- Quantum score matching with applications to learning thermal statesScore matching has driven major advances in classical generative learning by enabling models to learn from data without evaluating intractable normalization constants, or partition functions. Yet, extending this principle to quantum learning requires rethinking its foundations, as quantum states are
- When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit AssignmentReinforcement learning with verifiable rewards (RLVR) supervises mathematical reasoning through final-answer correctness, but provides little guidance on individual tokens. On-policy distillation (OPD) supplies dense feedback on student-generated responses, yet teacher preference need not reflect co
- ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid ControlHumanoid control, leveraging human demonstrations, has achieved diverse, agile, and natural locomotion behaviors through reinforcement learning (RL). While this paradigm has yielded remarkable performance in physical humanoid control, how to eliminate specific motions from learned policies remains i
- Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumerConsumers increasingly delegate purchasing decisions to Large Language Models (LLMs) acting as surrogate consumers. Using "Tool-Lab," an adaptation of information-board process tracing that places product attributes behind costly tool calls, we examine how marketing pricing cues (i.e., just-below pr
- AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving ScenariosVision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning. We introduce AnchorReasoning, a visually grounded reasoning dataset built
- LEAP-CBF: A Safety Filter for Uncertain Systems with Least-Effort Adversarial PotentialsControl barrier functions (CBF) are a popular safety filter to ensure safety for nonlinear dynamical systems. However, when the system is subject to uncertainties and disturbances, this requires the use of robust variants of CBFs, which can be difficult to construct and can be overly conservative, e
- MicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and InferenceMicroscaling quantization techniques are increasingly used to represent neural network parameters with 8 bits or fewer while preserving near-full precision accuracy. However, applying these methods efficiently in convolutional layers is not straightforward. A naive approach transfers full-precision
OpenRouter appstop 40 by tokens
- Hermes AgentHermes Agent is an open-source, self-improving AI agent by Nous Research that runs persistently with memory across sessions, and builds reus
- Kilo CodeKilo Code is an open-source AI coding agent that works across VS Code, JetBrains, and CLI to help developers ship code faster with agentic w
- ClineCline is an open-source AI coding agent that lives inside your IDE, autonomously exploring your codebase, editing files, running terminal co
- FreebuffPowerful coding models, funded by ads
- piThere are many coding agents, but this one is yours.
- omp
- ISEKAI ZEROAI adventures. Travel with your favorite characters
- Hello Minds, powered by Ethoswarm
- HighLevel
- FramerCreate a professional website with Framer’s no-code AI website builder. Design freely, manage CMS content, optimize SEO, collaborate, and pu
- ZCodeZCode combines the best AI agents with your existing tools so you can plan, code, review, and deploy without friction.
- Portkey AIControl panel for AI apps
- OpenClawOpenClaw is an open-source AI agent that connects to your messaging apps and takes real actions on your behalf, from running commands and br
- Nous Research APINous Research is an AI research team focused on advancing open-source AI.
- Claude CodeClaude Code is Anthropic's agentic coding tool that reads your entire codebase, plans and executes changes across files, runs tests, and ite
- DeepSeek Harness
- DescriptAI Video & Podcast Editor
- draco-cascade-bench
- Open WebUIOpen WebUI is a self-hosted AI platform providing a chat interface for Large Language Models.
- HackerAIA penetration testing agent that helps you find, validate, and report vulnerabilities with AI.
- CraftBuild and play AI RPGs
- Zoo Code
- Mira is the leading AI agent inside TelegramMira is a Telegram-native AI assistant.
- GDevelopAI assisted game engine.
- CodeGPT
- SillyTavernSillyTavern is the LLM frontend for power users, a chat interface that connects to any model and gives deep control through character creati
- LemonadeThe AI tool for Roblox games.
- websim
- SahasraAll-in-one creator store
- Command Code
- CodexA coding agent that helps you build and ship with AI
- Janitor AIJanitor AI is a chatbot platform where users create and chat with custom AI characters for interactive roleplay, storytelling, and immersive
- Roo CodeRoo Code is a VS Code extension that gives you a team of specialized AI agents with customizable modes for coding, architecture, debugging,
- ale-evolApplied research lab for frontier models
- MavenBio
- OpenHandsAI coding agent that can run commands, browse the web, call APIs
- Halluna
- Zed EditorAI code editor designed for high-performance collaboration
- pre.dev-agent
- Sophia's LoreBaryOrganize and enhance your roleplay creations.
YouTuberecent videos
- Claude Opus 5.5 AI: An Incredible Leap Forward
- Yes, Jev Is Insane, But There's A Catch
- DeepSeek’s Insane New Architecture
- Claude JUST found hidden DNA...
- AI labs may be hiding their biggest breakthroughs
- OpenAI JUST got HACKED...
- OpenAI's model just JAILBROKE ITSELF...
- Claude Opus 5.5 Just Did Something That Looks Like AGI
- Elon Musk Just Released Grok 4.7 — And It’s NOT What We Expected
- OpenAI Just Revealed Something Terrifying About Its AI Models
- This New AI Could Be the Biggest Breakthrough Since ChatGPT - Jev
- The most expensive 33 hours in WordPress history...
- An ex-OpenAI researcher just deleted language from the LLM...
- Did Google just kickstart the intelligence explosion?
Podcastslatest episodes
- 90 minutes of unfiltered product advice from Snap and Discord’s product chief | Peter Sellis
- How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)
- Why companies are becoming a series of loops | Anish Acharya (a16z)
- AI’s third era: the rise of persistent AI coworkers | Tara Seshan (Product Lead ChatGPT Work)
- How to close $100K+ enterprise deals, step by step | Jen Abel
- 🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)
- 🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science
- Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI
- Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC
- Humanity’s Last Invention — Richard Socher of Recursive
- A.I. Safety Goes Mainstream + a ‘Hard Fork’ Exit AMA
- The Ezra Klein Show: The A.I. Revolt Is Here
- What’s a Hard Fork?




