A newsletter that makes FOMO obsolete.

philosoraptor

if i fear missing out on reading this newsletter, did i win or lose


01 · The One Thing

The 90-minute war

Tuesday morning, Anthropic launched Claude Opus 5.5: Fable-5.1-level performance, 30% faster output, framed as beating GPT-5.6 Sol on CursorBench at one-third the cost. Ninety minutes later, OpenAI dropped GPT-6 Sol and Luna and cut API prices 50% on the spot.

This is the first model cycle where the launch-day discourse is about unit economics first and capability second. The timeline is wall-to-wall price-per-task memes. Polymarket is literally pricing “up to 93% cheaper coding costs.”

Signal: If you build on APIs, your cost basis got rewritten twice before lunch on Tuesday. Every agent product’s margins just moved. Recalculate build-vs-buy before you ship anything priced on old token math. Sources: techcrunch on OpenAI · macrumors on Anthropic

02 · Lab Watch

GPT-6 Sol + Luna, prices halved

OpenAI Sol for complex coding and agentic work, Luna for high-volume tasks. Sol now $2/$10 per 1M input/output tokens (was $4/$20), Luna $0.10/$0.50, cached reads 90% off. Permanent cuts, not promo.

Signal: OpenAI is passing inference-efficiency gains straight to developers. The Jevons playbook, live: cheaper thinking, more thinking.

Opus 5.5, and a November IPO looms

Anthropic List prices $4/$20 per 1M tokens (20% below Opus 5), but cache reads dropped 60% to $0.20/1M. Anthropic’s pitch: cache dominates agentic workloads, so typical costs fall ~40%. Sonnet 5.5 and Haiku 5.5 due in weeks. First release since Amodei’s “pace the frontier” call.

Signal: Both labs cut prices the same day via different mechanisms. The real benchmark is now cost per completed agent task, not cost per token.

Grok 4.7 ships, Copilot day one

xAI New largest base model, longer RL on multi-hour tasks, new safeguard stack, $2/$6 per 1M tokens. Rolling out to GitHub Copilot immediately. Independent Artificial Analysis testing matched xAI’s claims except Terminal-Bench 4.0: claimed 38.0%, measured 25.76%.

Signal: Vendor benchmarks remain guilty until proven innocent. Also: Copilot distribution gives xAI instant developer reach, which matters more than the benchmark gap.

Hardware blitz at Connect, plus a curtain pulled back

Meta Zuckerberg unveiled Muse Charm (keychain-sized pocket device for the Muse agent, December), Ray-Ban Meta Gen 3 at $449, camera-free audio glasses at $349, an FDA-cleared hearing-aid feature inside glasses, and Muse computer-use for Mac. Meanwhile Reuters found human contractors secretly making Muse’s phone calls on users’ behalf, one of whom made a racist reference on a call. Meta apologized and rolled the feature back. Separately: Muse is already straining Meta’s compute at ~700K DAU, and Amazon blocked Muse from shopping on its platform.

Signal: The “fully autonomous agent” still has humans behind the curtain, at Meta scale. And agents eating 10x projected tokens is a brutal unit-economics problem nobody has solved. Sources: techcrunch on Connect · gizmodo on contractors

Gemini 3.8 Flash TTS takes voice benchmarks

Google DeepMind New text-to-speech models: 100+ languages, 2,000 production voices, 30-second voice cloning, SynthID watermarking. Took #1 on Hume AI’s Voice Design Benchmark and Voice Arena preference evals.

Signal: Voice quality is now a leaderboard battleground, not a solved commodity. If your product has a voice, the frontier moved this week.

03 · Open Weights

NVIDIA Nemotron 3 Diarization

NVIDIA 100M-parameter speaker diarization (who-spoke-when, up to 8 speakers), streaming and offline, reportedly #1 on VoiceArena’s diarization board.

Signal: Tiny, practical, open. This quietly becomes infrastructure in every meeting-recording pipeline. See “Build on This” below.

Black Forest Labs FLUX 3 Action

Black Forest Labs 7B robotics world-action model predicting the next 32 robot actions from camera frames plus text. Claims 42.92% on NVIDIA’s RoboLab-120, beating NVIDIA’s own 16B Cosmos policy. FP8 variant runs on 24GB consumer GPUs. Commercial license terms unclear in coverage found.

Signal: Competitive robot policies on consumer GPUs could do for robotics what Stable Diffusion did for images. Verify the commercial license before building on it.

Alibaba Qwen-Image-2.1

Alibaba 7B open-weights text-to-image plus image editing, native 2K, day-zero support in Diffusers, ComfyUI, vLLM, SGLang.

Signal: Alibaba keeps shipping the most usable open image models. Day-zero ecosystem support means instant adoption.

Frontier watch

Moonshot’s Kimi K3 remains the first open model to top a board outright (#1, LMArena Frontend Code Arena). On Hugging Face, Qwen3.8-27B is now the most-liked model, with DeepSeek-V4-Pro and GLM-5.2 crowding the frontier band.

Signal: Open models now sit inside the frontier Elo band. The question shifted from “which model is best” to “which is best per dollar.”

04 · Money Flow

Island $400M at $6.4B · Tekever $580M at $6.4B

Same $6.4B valuation, same week, two theses: Island sells the enterprise browser as “scale AI agents securely” (Sequoia, Coatue, Insight in the round), Tekever sells AI drones weeks after a UK MoD surveillance contract worth up to £400M.

Signal: Money is flowing to AI with a government or enterprise-security buyer attached. Agent-security infra is commanding decacorn-adjacent valuations.

Ema raises $77M: “your SaaS apps become databases”

“AI employees” startup automating HR/IT/finance, led by Creaegis with Accel and Prosus. Valuation up 4x since 2024, $140M total raised.

Signal: The clearest articulation yet of the agent-eats-SaaS displacement thesis, now with revenue behind it. If you sell SaaS, this is your competitor’s pitch deck.

Go.AI raises $85M, profitable, on-prem

On-prem AI infra for regulated industries: 200+ customers, 8x ARR growth, 12.5M queries/day, profitable. Led by Updata Partners.

Signal: Sovereign AI is a real, profitable category. Banks, hospitals, and defense want frontier AI that never leaves the building.

Flatkey raises $10M: arbitraging the labs

Single API key across 100+ models and 1,000+ tools at 60-90% of official list prices. 10,000 developers in two months since July.

Signal: A startup growing this fast by undercutting list prices is itself a pricing signal. Margin compression is coming for undifferentiated model access.

05 · Leaderboard

Who’s king today

Claude Fable 5.1 holds the LMArena text crown at 1,514 Elo; Opus 5.5 debuted top-three on the Artificial Analysis Intelligence Index. Kimi K3 still the only open model to top a board outright. Seven of 86 vals.ai-evaluated models now clear 95% on SWE-Bench.

Signal: Benchmarks are saturating at the top, which is exactly why the labs pivoted the fight to price. Capability is table stakes; cost per task is the differentiator.

06 · The Argument

“Which one y’all taking”

AI Instagram and Threads are wall-to-wall with side-by-side price/performance charts from the 90-minute launch war: dueling AutomationBench and CursorBench claims, “up to 93% cheaper coding” memes, developers recalculating build-vs-buy in real time.

Signal: First cycle where economics lead capability in the discourse. When the crowd prices your infra like a commodity, build the thing on top, not the infra.

Altman and Amodei briefed the UN Security Council

Both CEOs jointly briefed the UN Security Council on AI risk on Sept 23, while Trump dismissed regulation talk at the General Assembly.

Signal: Mostly theater, but notable: the doomer-vs-accelerationist fight now has a diplomatic chapter. File under context, not action.

07 · Papers

Shopping by algorithm: how agentic AI deploys human heuristics as a surrogate consumer

Agents are becoming the buyers, and they shop with human-like heuristics.

Signal: If your checkout isn’t machine-readable, you’re invisible to the fastest-growing customer segment. The agentic-commerce wave is now a research topic, which means it’s about 12 months from being your problem.

Can LLMs reason about runtime behavior? A repository-level dynamic benchmark

Static benchmarks are saturating, so evaluation is moving to dynamic execution: can the model predict what code actually does when run?

Signal: Execution-based eval is the next frontier for coding agents. If you eval models, start tracking this line of work.

Agent-editing world model: rethinking world modeling for LLM agents

World models, but for LLM agents instead of robots: learning editable internal models of the environments agents act in.

Signal: Agents that simulate before acting beat agents that act and apologize. Worth a skim if you build multi-step agents.

Memory Attention

New attention variant aimed at long-horizon memory, landing the same week GitHub trending filled up with agent-memory repos.

Signal: Theory and practice converging on the same bottleneck: agents that can’t remember can’t compound. Read alongside hindsight and ai-memory below.

When and where to trust the teacher: unifying on-policy distillation and GRPO

Distillation theory catching up with practice: when teacher signals help, when they hurt, and how entropy mediates it.

Signal: Everyone distills now; few understand when it backfires. Practical if you train or fine-tune anything.

The theme: memory, skills, and harnesses

Nearly every fast mover is agent infrastructure. The gold rush isn’t models anymore; it’s the scaffolding around them.

  • vectorize-io/hindsight · +1,607 today · agent memory that learns
  • google/ax · +1,376 today · Google’s open agentic orchestration runtime
  • dream-num/univer · +1,060 today · office harness for agents
  • alibaba/open-code-review · +9,833 this week · biggest mover
  • affaan-m/ECC · +6,695 this week · harness perf optimization
  • stablyai/orca · +6,435 this week · fleet-of-agents IDE
  • Tencent/WeKnora · +4,522 this week · docs to queryable knowledge
  • addyosmani/agent-skills · +3,867 this week · production agent skills
  • JustVugg/colibri · +2,739 this week · frontier MoE in pure C
  • akitaonrails/ai-memory · +1,208 this week · long-term memory in Rust
  • obra/superpowers · +606 today · agentic skills framework
  • superdesigndev/treg · +470 today · “OpenRouter for agent tools”
  • strands-agents/harness-sdk · +463 today · build a harness end-to-end
  • HKUDS/CLI-Anything · +415 today · make all software agent-native

Signal: Three bets the crowd is making: agents need memory (hindsight, ai-memory), agents need skills (superpowers, agent-skills), agents need cheaper/faster harnesses (ECC, ax). Google open-sourcing orchestration plumbing is the tell that this layer is commoditizing fast.

09 · Hacker News

What is RLCD? The secret behind Jev

The training method behind TypeSafe’s Jev, explained. 51 points, early thread.

Signal: The technical companion to this week’s Jev discourse. Worth the read.

ArXiv receives multiyear commitments as an independent nonprofit

The commons gets funded. ArXiv secures its future as independent infrastructure. 286 points.

Signal: Quietly important: the paper pipeline you rely on just got durable funding.

Hackers influence ChatGPT and Gemini to direct users to scam centers

Prompt injection at population scale: poisoned content steering assistants toward scam operations. 114 points.

Signal: The attack surface of “agents that browse” is now a crime beat. If you ship a browsing agent, this is your threat model.

Also notable

Agents.md speaks Unix, and you should too: agent conventions as unix philosophy, small file, big idea. The Year of Internal Tools: the quiet enterprise AI boom is internal tooling, not products. Oracle invokes force majeure on New Mexico AI data center: the buildout is hitting physical limits. LinkedIn wins court order blocking mass scraping: training-data moats just got legal precedent.

10 · OpenRouter Apps

Where the tokens actually flow

OpenRouter Ranked by tokens routed through OpenRouter. Coding agents own this chart: Hermes Agent (Nous Research, ~49T), Kilo Code (~13T), Cline (~12T), Freebuff (~7T). The top of the market is agentic coding, and it’s not close.

Signal: Whatever the labs announce, the revealed preference of developers is coding agents. Build distribution where the tokens are.

New and climbing

DecodingTrust Agent Platform (jevagent) appeared Sept 22 and is already ranked: third-party eval infra for Jev-style models, live within days of launch. Also new: sel-jev-rubric-judge (Sept 20), rustllm (Sept 18), LuckyRobots worker fleet (Sept 18).

Signal: Jev got an eval-harness ecosystem within 72 hours of launch. The Jevons paradox in real time: collapse the cost of thinking and new apps sprout before the paint dries.

11 · Watch

Yes, Jev Is Insane, But There’s A Catch

Two Minute Papers Two Minute Papers · Sep 22

Signal: The clearest 10-minute technical take on Jev’s trade-offs.

Claude Opus 5.5 AI: An Incredible Leap Forward

Two Minute Papers Two Minute Papers · Sep 24

An ex-OpenAI researcher just deleted language from the LLM…

Fireship Fireship · Sep 21

Signal: That’s the TypeSafe story in Fireship’s 5-minute format: giving up string generation entirely. Good shareable explainer.

AI labs may be hiding their biggest breakthroughs

Wes Roth Wes Roth · Sep 22

Elon Musk Just Released Grok 4.7 — And It’s NOT What We Expected

The AI Grid The AI Grid · Sep 22

12 · Listen

Jev: System One models for Prod, not God — with Diogo Almeida

Latent Space Latent Space · Sep 21 · episode

Signal: The founder interview behind the model.

Bio-security is an AI Arms Race

Latent Space Latent Space · Sep 23 · episode

90 minutes of unfiltered product advice from Snap and Discord’s product leaders

Lenny’s Podcast Lenny’s Podcast · Sep 20 · episode

13 · Read

Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?

Lenny’s Newsletter Lenny’s Newsletter · Sep 22 · read

Signal: The launch war, adjudicated blind. The most useful single comparison of the two models.

I left Claude for months. Opus 5.5 is why I’m back

Lenny’s Newsletter · Sep 22 · read

TLDR AI — Sep 24 edition: Gemini TTS, Claude’s novel enzyme, Google private memory

TLDR TLDR AI · Sep 24 · read

14 · Safe to Ignore

Permission to skip

Skip: The UN Security Council AI briefing. Diplomacy is not a product signal. Nothing announced, nothing decided. Skip: Bessemer’s $5.75B in new funds. Dry-powder press releases don’t change your roadmap. Note the direction of travel, then move on. Skip: Grok 4.7’s Terminal-Bench gap (38% claimed, 25.8% measured). Unless you benchmark models for a living, file under “vendor charts lie a little” and move on. Skip: F-Droid 2.0 and the Nokia Design Archive. Lovely threads. Not FOMO. Read them on the weekend.

15 · Build on This

The meeting-intel refinery

NVIDIA just open-sourced a 100M-parameter diarization model that does who-spoke-when and runs anywhere. Luna-class models now cost $0.10 per million input tokens. The two hardest parts of “upload a meeting, get intelligence out” just went to ~zero.

Product: drop in a recording, get chapters, action items, and a searchable transcript. Sell it as a $29 one-time credit pack for 10 hours of audio. The wedge: every incumbent charges per seat per month; this charges per meeting, once.

16 · Appendix

Everything below didn’t make the cut. Tell me what I promoted wrongly or buried unfairly and issue #2 gets sharper.

Hacker Newstop 30
  1. F-Droid 2.0
    346 pts · 92 comments · by daveoc64 · 09-24 08:26 · discuss
  2. Show HN: Whiteboard (YC W26) – An open-source IDE for thoughtful software design
    15 pts · 0 comments · by sidharthkmenon · 09-24 10:21 · discuss
  3. Two-tier encryption in the UK
    260 pts · 252 comments · by ReturnoftheHack · 09-24 03:39 · discuss
  4. Agents.md speaks Unix, and you should too
    10 pts · 1 comments · by fmind-dev · 09-24 09:31 · discuss
  5. Nokia Design Archive (2025)
    175 pts · 92 comments · by pillars · 09-24 02:49 · discuss
  6. WaveDigger: Dig into wireless signals to discover their physical locations
    25 pts · 3 comments · by 882542F3884314B · 09-23 05:57 · discuss
  7. GitHub has not removed malicious imitation software after 3 weeks
    115 pts · 44 comments · by hermitcrab · 09-24 08:50 · discuss
  8. B5-BJ2 – Ice Cream Barges – Concrete Ship Constructors (2023)
    8 pts · 3 comments · by nixass · 09-21 15:31 · discuss
  9. Linux support is coming to Snapdragon X2 series
    574 pts · 241 comments · by aaronday · 09-23 15:38 · discuss
  10. Experiencing writing at our recent Chinese calligraphy workshop
    7 pts · 1 comments · by surprisetalk · 09-24 08:40 · discuss
  11. The science of Monkey Island: can grog dissolve a metal mug that fast?
    91 pts · 16 comments · by zdw · 09-23 08:55 · discuss
  12. Ideas on modernizing the open-source desktop
    330 pts · 411 comments · by signa11 · 09-23 19:52 · discuss
  13. Federal judge orders Texas to air condition all prisons by the end of 2029
    58 pts · 43 comments · by bonefishgrill · 09-24 09:15 · discuss
  14. RAM: the forgotten history (2024)
    96 pts · 3 comments · by Luc · 09-22 08:12 · discuss
  15. When the Debugger Lies
    54 pts · 17 comments · by hasheddan · 09-22 04:02 · discuss
  16. ArXiv receives multiyear commitments to support it as an independent nonprofit
    286 pts · 38 comments · by JohnHammersley · 09-23 15:45 · discuss
  17. Enjoy Every Sandwich
    152 pts · 68 comments · by NaOH · 09-22 11:19 · discuss
  18. Coulomb's law remains tricky to test at home
    16 pts · 15 comments · by surprisetalk · 09-22 06:00 · discuss
  19. The newest ESP32 can run Linux and it's getting close to a Raspberry Pi
    170 pts · 76 comments · by adunk · 09-24 04:08 · discuss
  20. VSCode's SSH Agent Is Bananas (2025)
    293 pts · 190 comments · by Rapzid · 09-23 14:01 · discuss
  21. What Is RLCD? The Secret Behind Jev
    51 pts · 6 comments · by tnspacetime · 09-24 05:21 · discuss
  22. Contrastive Language Models
    145 pts · 42 comments · by erichocean · 09-23 21:20 · discuss
  23. The "Windows XP Box" (2003)
    214 pts · 44 comments · by doubletwoyou · 09-21 20:09 · discuss
  24. Oracle invokes force majeure on New Mexico AI data center
    18 pts · 2 comments · by dgellow · 09-24 08:58 · discuss
  25. The Year of Internal Tools
    54 pts · 19 comments · by thecodemonkey · 09-24 00:23 · discuss
  26. Fixing the Portobello Police Station Clock
    510 pts · 112 comments · by avidly · 09-23 08:18 · discuss
  27. Hackers influence ChatGPT and Gemini to direct users to scam centers
    114 pts · 38 comments · by ArielSimon · 09-24 04:54 · discuss
  28. LinkedIn wins court order blocking mass scraping of user data
    40 pts · 27 comments · by ilamont · 09-24 09:02 · discuss
  29. Why 'What's Opera, Doc?' looks like that
    125 pts · 22 comments · by CharlesW · 09-23 08:37 · discuss
  30. Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering
    100 pts · 32 comments · by phatak-dev · 09-24 07:33 · discuss
GitHub trending · daily14 repos
  1. rohitg00/ai-engineering-from-scratch
    Learn it. Build it. Ship it for others.
    Python · +310 stars today
  2. vectorize-io/hindsight
    Hindsight: Agent Memory That Learns
    Python · +1,607 stars today
  3. dream-num/univer
    The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.
    TypeScript · +1,060 stars today
  4. google/ax
    Google's open agentic orchestration runtime
    Go · +1,376 stars today
  5. NVIDIA/Model-Optimizer
    A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
    Python · +22 stars today
  6. FxEmbed/FxEmbed
    Fix X/Twitter and Bluesky embeds! Use multiple images, videos, polls, translations and more on Discord, Telegram and others
    TypeScript · +165 stars today
  7. anthropics/financial-services
    Python · +510 stars today
  8. HKUDS/CLI-Anything
    "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
    Python · +415 stars today
  9. mvt-project/mvt
    MVT (Mobile Verification Toolkit) helps with conducting forensics of mobile devices in order to find signs of a potential compromise.
    Python · +275 stars today
  10. obra/superpowers
    An agentic skills framework & software development methodology that works.
    Shell · +606 stars today
  11. strands-agents/harness-sdk
    Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.
    Python · +463 stars today
  12. julyx10/lap
    An offline-first photo manager for large local libraries
    Vue · +151 stars today
  13. superdesigndev/treg
    OpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn
    Python · +470 stars today
  14. leejet/stable-diffusion.cpp
    Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
    C++ · +69 stars today
GitHub trending · weekly20 repos
  1. anthropics/financial-services
    Python · +1,895 stars this week
  2. alibaba/open-code-review
    Secure, fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.
    Go · +9,833 stars this week
  3. anthropics/claude-code
    Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
    TypeScript · +2,762 stars this week
  4. affaan-m/ECC
    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
    JavaScript · +6,695 stars this week
  5. Tencent/WeKnora
    Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
    Go · +4,522 stars this week
  6. addyosmani/agent-skills
    Production-grade engineering skills for AI coding agents.
    JavaScript · +3,867 stars this week
  7. anthropics/knowledge-work-plugins
    Open source repository of plugins primarily intended for knowledge workers to use in Claude Cowork
    Python · +1,358 stars this week
  8. stablyai/orca
    Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
    TypeScript · +6,435 stars this week
  9. TencentCloud/Octop
    A smarter, self-hosted AI assistant — multi-user, multi-agent.
    Python · +1,877 stars this week
  10. davila7/claude-code-templates
    CLI tool for configuring and monitoring Claude Code
    Python · +613 stars this week
  11. JustVugg/colibri
    Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
    C · +2,739 stars this week
  12. vectorize-io/hindsight
    Hindsight: Agent Memory That Learns
    Python · +3,181 stars this week
  13. cloudflare/quiche
    🥧 Savoury implementation of the QUIC transport protocol and HTTP/3
    Rust · +707 stars this week
  14. Fission-AI/OpenSpec
    Spec-driven development (SDD) for AI coding assistants.
    TypeScript · +1,538 stars this week
  15. akitaonrails/ai-memory
    Solution for long term memory for agent coding CLIs and to facilitate handoff between different agent vendors
    Rust · +1,208 stars this week
  16. pytorch/pytorch
    Tensors and Dynamic neural networks in Python with strong GPU acceleration
    Python · +196 stars this week
  17. cline/cline
    Autonomous coding agent as an SDK, IDE extension, or CLI assistant.
    TypeScript · +1,177 stars this week
  18. LibreChat-AI/LibreChat
    Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API, Azure, Groq, o1, GPT-5, Mistral, OpenRouter, Vertex AI, Gemini, Artifacts, AI model switching, message search, Code Interpreter, langchain, DALL-E-3, OpenAPI Actions, Functions, Secure Multi-User Auth, Presets, open-source for self-hosting. Active
    TypeScript · +949 stars this week
  19. superdesigndev/treg
    OpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn
    Python · +1,067 stars this week
  20. cilium/cilium
    eBPF-based Networking, Security, and Observability
    Go · +447 stars this week
Hugging Facemost liked · 30
  1. Qwen/Qwen3.8-27B
    16,201 likes · 6,765,008 downloads · image-text-to-text
  2. black-forest-labs/FLUX.1-dev
    14,974 likes · 710,873 downloads · text-to-image
  3. deepseek-ai/DeepSeek-R1
    14,287 likes · 822,578 downloads · text-generation
  4. moonshotai/Kimi-K3
    11,504 likes · 1,774,987 downloads · image-text-to-text
  5. stabilityai/stable-diffusion-xl-base-1.0
    8,224 likes · 3,506,330 downloads · text-to-image
  6. meta-llama/Llama-3.1-8B-Instruct
    7,876 likes · 6,247,724 downloads · text-generation
  7. CompVis/stable-diffusion-v1-4
    7,096 likes · 658,816 downloads · text-to-image
  8. hexgrad/Kokoro-82M
    6,996 likes · 11,870,208 downloads · text-to-speech
  9. meta-llama/Meta-Llama-3-8B
    6,706 likes · 255,181 downloads · text-generation
  10. openai/whisper-large-v3
    6,376 likes · 4,700,322 downloads · automatic-speech-recognition
  11. sentence-transformers/all-MiniLM-L6-v2
    6,119 likes · 250,598,416 downloads · sentence-similarity
  12. black-forest-labs/FLUX.1-schnell
    5,997 likes · 520,328 downloads · text-to-image
  13. Qwen/Qwen3.8-Flash-Next
    5,660 likes · 830,208 downloads · image-text-to-text
  14. MiniMaxAI/MiniMax-H3
    5,647 likes · 3,640,535 downloads · image-text-to-video
  15. deepseek-ai/DeepSeek-V4-Pro
    5,589 likes · 514,445 downloads · text-generation
  16. Tongyi-MAI/Z-Image-Turbo
    5,354 likes · 622,744 downloads · text-to-image
  17. openai/gpt-oss-120b
    5,296 likes · 4,767,451 downloads · text-generation
  18. zai-org/GLM-5.2
    5,132 likes · 927,677 downloads · text-generation
  19. meta-llama/Meta-Llama-3-8B-Instruct
    5,104 likes · 1,183,533 downloads · text-generation
  20. openai/gpt-oss-20b
    5,093 likes · 6,692,738 downloads · text-generation
  21. bigscience/bloom
    5,070 likes · 15,034 downloads · text-generation
  22. stabilityai/stable-diffusion-3-medium
    5,052 likes · 3,751 downloads · text-to-image
  23. Lightricks/LTX-2.5
    4,974 likes · 1,637,601 downloads · image-to-video
  24. meta-llama/Llama-2-7b-chat-hf
    4,875 likes · 504,057 downloads · text-generation
  25. mistralai/Mixtral-8x7B-Instruct-v0.1
    4,756 likes · 215,760 downloads
  26. unsloth/Qwen3.8-27B-GGUF
    4,577 likes · 7,063,930 downloads
  27. meta-llama/Llama-2-7b
    4,558 likes · 115 downloads · text-generation
  28. baidu/Unlimited-OCR
    4,284 likes · 2,061,101 downloads · image-text-to-text
  29. deepseek-ai/DeepSeek-V3
    4,251 likes · 1,099,848 downloads · text-generation
  30. mistralai/Mistral-7B-v0.1
    4,184 likes · 438,368 downloads · text-generation
arXivcs.AI / cs.CL / cs.LG · newest 25
  1. On the Diffusibility of High-Dimensional Latents
    Chao Feng, Zhiyang Xu, Bowei Chen, Yuanjun Xiong · 2026-09-23
    Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reco
  2. Contrastive Learning for Authorship Verification
    Peter Kirby · 2026-09-23
    Our results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context length, and random text span data augmentation as important factors
  3. StudentBench: AI and human tutoring yield equivalent GRE learning gains
    Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi · 2026-09-23
    Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection
  4. Where Should I Join? Robot Group Joining via Language-Guided Goal Prediction
    Zilin Fang, Zishuo Wang, Gim Hee Lee, David Hsu · 2026-09-23
    Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity and formation. This is a highly semantic task, yet an important capability for ap
  5. Even Sharper Bounds for Transductive Learning and Its Applications
    Yingzhen Yang · 2026-09-23
    We introduce Sharper Transductive Local Complexity (STLC), a localized complexity method for transductive learning under uniform sampling without replacement. The construction starts from a Bernstein-type concentration inequality for the supremum of the test--train empirical process. Its proof uses
  6. Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark
    Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala · 2026-09-23
    Large language models (LLMs) are increasingly used in coding tasks, but their ability to reason about code execution remains unclear. Existing repository-level QA benchmarks mainly evaluate static code understanding and often rely on LLM-based evaluation, while execution-reasoning benchmarks are mos
  7. Nonequilibrium Phases of Repulsive Self-Attention: Chaos, Attention Condensation, and Emergent Locality
    Qucheng Gao, Zuyi Yang, Xiao Chen · 2026-09-23
    We study the nonequilibrium dynamics of a minimal recurrent transformer with $N$ normalized tokens, $Q=K=I$, and a negative value map $V=-I$. Similarity-based attention selects nearby representations, while the negative value map drives tokens away from the selected field. This feedback can continua
  8. Order-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning
    Zhixu Silvia Tao · 2026-09-23
    Reordering a set of mathematical rules without changing its meaning should preserve the correct answer, but must a model's internal representations stay invariant too? We investigate this question using synthetic multi-step function-composition problems, each presented under multiple rule orderings
  9. Minimal-Norm Univariate Two-Layer ReLU Classification: Exact Solutions and Global Optimality with Skip Connections
    Karolina Drabik, Ben Lewis, Antoni Puch, Etienne Boursier · 2026-09-23
    We study minimal-norm interpolation and $\ell_2$-regularized logistic-loss minimization for binary classification by univariate two-layer ReLU networks. We give complete geometric characterizations of the optimal classifiers in function space, resolving how the solutions depend on whether hidden-lay
  10. Cross-Scale Transfer Learning for Depression Severity Prediction: From PHQ-8 to HAMD-17 Across Languages and Clinical Paradigms
    Wenjie Feng, Sahba Zojaji, Satoshi Nakamura · 2026-09-23
    This work addresses continuous depression-severity score prediction from clinical interview transcripts under data scarcity. We propose a sequential low-rank adaptation (LoRA) protocol for cross-scale transfer: a Qwen3 backbone with a bounded regression head is first fine-tuned on the English DAIC-W
  11. Context-Continuous Preference Learning for Exoskeleton Personalization
    Sunin Baek, Sungwoo Park, Daekyum Kim · 2026-09-23
    Personalizing exoskeleton assistance across operating conditions is constrained by the time and physical effort required to collect user feedback. We examined whether a user's preference landscape varies smoothly across operating conditions and when this continuity supports learning from limited fee
  12. Repairability of Inexact Solvers in Recursive State Estimation with Machine Learning
    Yanjun Ji, Dennis Willsch, Orkun Şensebat, Priyanka Arkalgud Ganeshamurthy · 2026-09-23
    Recursive state estimation often executes approximate numerical solutions inside a feedback loop, where highly accurate local steps do not guarantee better overall results. For a fixed linear Kalman model, we characterize when a correction within a prescribed subspace and norm budget can meet a loca
  13. Agent-Editing World Model: Rethinking World Modeling for LLM Agents
    Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng · 2026-09-23
    Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool res
  14. Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model
    Xiwen Chen, Rigaudiere Z. Li, Zhiruo Zhou, Xiaojun Zhu · 2026-09-23
    Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow never moves the manipulated object; retraining it with latent-only losses only trades stillness for te
  15. Learning Holographic Reduced Representations with Clifford Variational Autoencoders
    Mohamed Malek Abid, P. Michael Furlong · 2026-09-23
    Vector Symbolic Algebras project data structures into a hyperdimensional vector space through the application of their vector algebras to randomly generated atomic vector symbols and fractional power encodings of real-valued data. Embedding unstructured data remains an open question. We present \tex
  16. Learning Collective Dynamics with Differentiable Gaussian Representations
    Jianxiang Ma, Mingfu Zhang, Xiaocui Yang, Yichen Gao · 2026-09-23
    Collective responses depend on individual differences, contact opportunities, and accumulated experience. Learning their dynamics from aggregate counts requires connecting a population's response distribution to both current observations and future behavior. We introduce Differentiable Gaussian Dyna
  17. Memory Attention
    Jiale Kang · 2026-09-23
    Language models typically construct attention values from contextual hidden states, even when some of their content may be reusable across contexts. We investigate whether token-indexed memory can replace the dedicated value projection when complemented by contextual information. We propose Memory A
  18. Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following
    Niklas Scholz, David Thulke, Abdallah Nasir, Will Allred · 2026-09-23
    Fine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings transfer to machine translation (MT) fine-tuning and to MT-specific inst
  19. Quantum score matching with applications to learning thermal states
    Yulong Dong, Jiaqi Leng · 2026-09-23
    Score matching has driven major advances in classical generative learning by enabling models to learn from data without evaluating intractable normalization constants, or partition functions. Yet, extending this principle to quantum learning requires rethinking its foundations, as quantum states are
  20. When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment
    Jie Zhang, Jingxiao Yang, Zhehao Huang, Yuhang Liu · 2026-09-23
    Reinforcement learning with verifiable rewards (RLVR) supervises mathematical reasoning through final-answer correctness, but provides little guidance on individual tokens. On-policy distillation (OPD) supplies dense feedback on student-generated responses, yet teacher preference need not reflect co
  21. ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control
    Xukun Luan, Zhongxiang Lei, Chen Gong, Shaowei Li · 2026-09-23
    Humanoid control, leveraging human demonstrations, has achieved diverse, agile, and natural locomotion behaviors through reinforcement learning (RL). While this paradigm has yielded remarkable performance in physical humanoid control, how to eliminate specific motions from learned policies remains i
  22. Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer
    Davood Wadi, Yu Ma · 2026-09-23
    Consumers increasingly delegate purchasing decisions to Large Language Models (LLMs) acting as surrogate consumers. Using "Tool-Lab," an adaptation of information-board process tracing that places product attributes behind costly tool calls, we examine how marketing pricing cues (i.e., just-below pr
  23. AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios
    Zhipeng Bao, Wenjie Zhao, Tianle Zhu, Haohua Que · 2026-09-23
    Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning. We introduce AnchorReasoning, a visually grounded reasoning dataset built
  24. LEAP-CBF: A Safety Filter for Uncertain Systems with Least-Effort Adversarial Potentials
    Oswin So, Eric Yu, Chuchu Fan · 2026-09-23
    Control barrier functions (CBF) are a popular safety filter to ensure safety for nonlinear dynamical systems. However, when the system is subject to uncertainties and disturbances, this requires the use of robust variants of CBFs, which can be difficult to construct and can be overly conservative, e
  25. MicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and Inference
    Romain Facq, Sami Ben Ali, Olivier Sentieys · 2026-09-23
    Microscaling quantization techniques are increasingly used to represent neural network parameters with 8 bits or fewer while preserving near-full precision accuracy. However, applying these methods efficiently in convolutional layers is not straightforward. A naive approach transfers full-precision
OpenRouter appstop 40 by tokens
  1. Hermes Agent
    49.10T tokens · created 2026-03-12
    Hermes Agent is an open-source, self-improving AI agent by Nous Research that runs persistently with memory across sessions, and builds reus
  2. Kilo Code
    13.19T tokens · created 2025-04-09
    Kilo Code is an open-source AI coding agent that works across VS Code, JetBrains, and CLI to help developers ship code faster with agentic w
  3. Cline
    11.56T tokens · created 2024-10-09
    Cline is an open-source AI coding agent that lives inside your IDE, autonomously exploring your codebase, editing files, running terminal co
  4. Freebuff
    6.72T tokens · created 2026-06-18
    Powerful coding models, funded by ads
  5. pi
    2.56T tokens · created 2026-02-15
    There are many coding agents, but this one is yours.
  6. omp
    1.98T tokens · created 2026-05-15
  7. ISEKAI ZERO
    1.59T tokens · created 2026-01-21
    AI adventures. Travel with your favorite characters
  8. Hello Minds, powered by Ethoswarm
    1.50T tokens · created 2026-06-18
  9. HighLevel
    1.47T tokens · created 2026-07-31
  10. Framer
    1.45T tokens · created 2026-04-18
    Create a professional website with Framer’s no-code AI website builder. Design freely, manage CMS content, optimize SEO, collaborate, and pu
  11. ZCode
    1.42T tokens · created 2026-02-12
    ZCode combines the best AI agents with your existing tools so you can plan, code, review, and deploy without friction.
  12. Portkey AI
    1.33T tokens · created 2023-12-18
    Control panel for AI apps
  13. OpenClaw
    1.29T tokens · created 2026-01-30
    OpenClaw is an open-source AI agent that connects to your messaging apps and takes real actions on your behalf, from running commands and br
  14. Nous Research API
    1.23T tokens · created 2026-02-26
    Nous Research is an AI research team focused on advancing open-source AI.
  15. Claude Code
    1.16T tokens · created 2025-12-23
    Claude Code is Anthropic's agentic coding tool that reads your entire codebase, plans and executes changes across files, runs tests, and ite
  16. DeepSeek Harness
    1.11T tokens · created 2026-08-14
  17. Descript
    1.08T tokens · created 2025-06-12
    AI Video & Podcast Editor
  18. draco-cascade-bench
    0.87T tokens · created 2026-06-17
  19. Open WebUI
    0.73T tokens · created 2024-07-28
    Open WebUI is a self-hosted AI platform providing a chat interface for Large Language Models.
  20. HackerAI
    0.58T tokens · created 2024-02-13
    A penetration testing agent that helps you find, validate, and report vulnerabilities with AI.
  21. Craft
    0.49T tokens · created 2026-05-25
    Build and play AI RPGs
  22. Zoo Code
    0.46T tokens · created 2026-05-29
  23. Mira is the leading AI agent inside Telegram
    0.45T tokens · created 2026-04-13
    Mira is a Telegram-native AI assistant.
  24. GDevelop
    0.39T tokens · created 2025-06-02
    AI assisted game engine.
  25. CodeGPT
    0.37T tokens · created 2026-08-21
  26. SillyTavern
    0.37T tokens · created 2023-07-21
    SillyTavern is the LLM frontend for power users, a chat interface that connects to any model and gives deep control through character creati
  27. Lemonade
    0.31T tokens · created 2025-12-28
    The AI tool for Roblox games.
  28. websim
    0.30T tokens · created 2025-05-14
  29. Sahasra
    0.28T tokens · created 2025-07-22
    All-in-one creator store
  30. Command Code
    0.27T tokens · created 2026-05-29
  31. Codex
    0.23T tokens · created 2026-01-10
    A coding agent that helps you build and ship with AI
  32. Janitor AI
    0.22T tokens · created 2023-12-09
    Janitor AI is a chatbot platform where users create and chat with custom AI characters for interactive roleplay, storytelling, and immersive
  33. Roo Code
    0.22T tokens · created 2024-12-09
    Roo Code is a VS Code extension that gives you a team of specialized AI agents with customizable modes for coding, architecture, debugging,
  34. ale-evol
    0.22T tokens · created 2025-05-05
    Applied research lab for frontier models
  35. MavenBio
    0.21T tokens · created 2026-06-08
  36. OpenHands
    0.21T tokens · created 2024-09-24
    AI coding agent that can run commands, browse the web, call APIs
  37. Halluna
    0.21T tokens · created 2026-06-12
  38. Zed Editor
    0.19T tokens · created 2025-04-14
    AI code editor designed for high-performance collaboration
  39. pre.dev-agent
    0.19T tokens · created 2026-06-01
  40. Sophia's LoreBary
    0.18T tokens · created 2025-08-27
    Organize and enhance your roleplay creations.
YouTuberecent videos
  1. Claude Opus 5.5 AI: An Incredible Leap Forward
    TwoMinutePapers · 2026-09-24
  2. Yes, Jev Is Insane, But There's A Catch
    TwoMinutePapers · 2026-09-22
  3. DeepSeek’s Insane New Architecture
    TwoMinutePapers · 2026-09-18
  4. Claude JUST found hidden DNA...
    WesRoth · 2026-09-24
  5. AI labs may be hiding their biggest breakthroughs
    WesRoth · 2026-09-22
  6. OpenAI JUST got HACKED...
    WesRoth · 2026-09-19
  7. OpenAI's model just JAILBROKE ITSELF...
    WesRoth · 2026-09-18
  8. Claude Opus 5.5 Just Did Something That Looks Like AGI
    TheAIGRID · 2026-09-24
  9. Elon Musk Just Released Grok 4.7 — And It’s NOT What We Expected
    TheAIGRID · 2026-09-22
  10. OpenAI Just Revealed Something Terrifying About Its AI Models
    TheAIGRID · 2026-09-21
  11. This New AI Could Be the Biggest Breakthrough Since ChatGPT - Jev
    TheAIGRID · 2026-09-18
  12. The most expensive 33 hours in WordPress history...
    Fireship · 2026-09-23
  13. An ex-OpenAI researcher just deleted language from the LLM...
    Fireship · 2026-09-21
  14. Did Google just kickstart the intelligence explosion?
    Fireship · 2026-09-17
Podcastslatest episodes
  1. 90 minutes of unfiltered product advice from Snap and Discord’s product chief | Peter Sellis
    Lenny's Podcast · Sun, 20 Sep 2026
  2. How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)
    Lenny's Podcast · Tue, 08 Sep 2026
  3. Why companies are becoming a series of loops | Anish Acharya (a16z)
    Lenny's Podcast · Sun, 06 Sep 2026
  4. AI’s third era: the rise of persistent AI coworkers | Tara Seshan (Product Lead ChatGPT Work)
    Lenny's Podcast · Sun, 30 Aug 2026
  5. How to close $100K+ enterprise deals, step by step | Jen Abel
    Lenny's Podcast · Sun, 23 Aug 2026
  6. 🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics)
    Latent Space · Wed, 23 Sep 2026
  7. 🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science
    Latent Space · Tue, 22 Sep 2026
  8. Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI
    Latent Space · Mon, 21 Sep 2026
  9. Underwriting Superintelligence: Backing Agents you can Sue — Rune Kvist, AIUC
    Latent Space · Wed, 16 Sep 2026
  10. Humanity’s Last Invention — Richard Socher of Recursive
    Latent Space · Mon, 14 Sep 2026
  11. A.I. Safety Goes Mainstream + a ‘Hard Fork’ Exit AMA
    Hard Fork · Fri, 18 Sep 2026
  12. The Ezra Klein Show: The A.I. Revolt Is Here
    Hard Fork · Fri, 11 Sep 2026
  13. What’s a Hard Fork?
    Hard Fork · Tue, 27 Sep 2022