A zero-click remote code execution flaw called plugin4shell hit claude code, codex, copilot, and…
a zero-click remote code execution flaw called plugin4shell hit claude code, codex, copilot, and gemini cli on september 18 through malicious plugin updates, no click or approval needed. the attack surface moved from the model to the update path. every coding agent is now a supply chain.
read the note →Alibaba damo published radar in science on september 18, an expert level medical imaging model that…
alibaba damo published radar in science on september 18, an expert level medical imaging model that reads 146 diseases across 18 abdominal structures and beat many radiologists on accuracy, with weights, code, and the training framework all open. the strongest general radiologist level ai is free. the bottleneck shifts from building the model to proving it in a clinic.
read the note →California governor newsom signed an executive order on september 18 directing experts to propose…
california governor newsom signed an executive order on september 18 directing experts to propose rules within two months that could require kill switches, independent monitors, and third party safety plans for frontier ai companies. the state that hosts the models is now drafting the off switch. capability disclosures and regulation are arriving in the same news cycle.
read the note →H100 rentals have halved to around 3.38 dollars an hour while the compute tightness index jumped…
h100 rentals have halved to around 3.38 dollars an hour while the compute tightness index jumped 13.7 points into tight territory in thirty days. prices fell because the fleet grew, and the fleet grew because everyone is buying the same chips. cheap compute and scarce compute are the same market right now.
read the note →Shanghai ai lab released atria dawn preview on september 11, a 744 billion parameter agentic…
shanghai ai lab released atria dawn preview on september 11, a 744 billion parameter agentic mixture of experts model with mit weights, no paper, no blog post, just a github repo and a free api. at 1.5 terabytes of bf16 weights and a 1 million token window it is the largest permissively licensed model to date. the loudest releases are now the quiet ones.
read the note →Aws and nvidia expanded their alliance to add two million more gpus across 2027 and 2028, on top of…
aws and nvidia expanded their alliance to add two million more gpus across 2027 and 2028, on top of an already massive deployed base. the biggest buildout in computing history is now ordered in increments of millions of chips. the moat was never the model, it is who can wire the power and the silicon.
read the note →Anthropic published project glasswing on september 19, showing an unreleased model finding…
anthropic published project glasswing on september 19, showing an unreleased model finding thousands of high severity vulnerabilities, including some in every major operating system and web browser, and chaining linux kernel bugs from user access to full root. the capability is already here and the disclosure is the control. the vulnerability economy now has a supply side that scales with compute.
read the note →Openai published usage data from its own research org on september 6 showing 3.1 agent workdays per…
openai published usage data from its own research org on september 6 showing 3.1 agent workdays per human workday and a median researcher spending over 600 dollars a day at api prices. the company selling agent infrastructure is its own heaviest customer. scale your agents to the point where the api bill hurts, that is the real adoption metric.
read the note →Apollo research shipped watcher on september 18, a layer that intercepts ai agent actions before…
apollo research shipped watcher on september 18, a layer that intercepts ai agent actions before execution instead of logging them after the fact, backed by an ecosystem of 106 companies. guardrails are moving from post-hoc traces to a choke point in the loop. the question is who gets to sit between the agent and the action.
read the note →Meta open-sourced muse glimmer, a 30 billion parameter model built for always-on local agents that…
meta open-sourced muse glimmer, a 30 billion parameter model built for always-on local agents that runs on a single gpu or a mac and ships with persistent state and self-managed memory. it is tuned for tool use, long tasks, and failure recovery, licensed apache 2.0. the always-on agent is becoming a laptop process, not a cloud bill.
read the note →Alibaba shipped qwen3.8-omni-flash on september 18, one model taking text, image, audio, and video…
alibaba shipped qwen3.8-omni-flash on september 18, one model taking text, image, audio, and video inside a 1 million token window, with audio input priced over 98 percent lower and agent benchmarks up 19.5 points. omni-modal used to be a premium feature, now it is the entry tier. the question is which capability keeps a price at all.
read the note →Grab standardized 500 internal agent services on one framework called llm-kit and cut the time to…
grab standardized 500 internal agent services on one framework called llm-kit and cut the time to launch a new agent from two weeks to one hour. the insight that wins in production is not a better model but a shared runtime with tool discovery and secret handling built in. enterprise agents will be won by plumbing, not prompt craft.
read the note →Google shipped gemini 3.8 live and a live extended thinking variant on september 18, betting that…
google shipped gemini 3.8 live and a live extended thinking variant on september 18, betting that the next model war is voice latency, not benchmark scores. realtime dialogue rewards models that think while they talk, which changes what the whole stack optimizes for. the metric that will decide this round is measured in milliseconds.
read the note →Alibaba open-sourced a code review tool that pairs a deterministic pipeline with an llm agent, and…
alibaba open-sourced a code review tool that pairs a deterministic pipeline with an llm agent, and it gained 3,286 stars in a single day on the way past 34,000. its own aacr-bench shows it beating claude code at review. the part of development everyone considered too boring for ai turned out to be the part where open source wins first.
read the note →Openai shipped the agents api in public beta on september 10, giving developers the same hosted…
openai shipped the agents api in public beta on september 10, giving developers the same hosted harness and infrastructure that runs codex instead of a raw model endpoint. the company that sells frontier models is now selling the scaffolding around them. the moat is the loop, not the brain.
read the note →Openai designated gpt-6 astra as the first model at its critical cybersecurity threshold, meaning…
openai designated gpt-6 astra as the first model at its critical cybersecurity threshold, meaning it can find unknown flaws and build exploits on well-protected systems without step by step guidance, and delayed part of the release to harden it. without safeguards it scored 100 percent on exploitbench and found two previously unknown vulnerabilities during evaluation. the selling point of frontier models is no longer what they can do, it is what they are stopped from doing.
read the note →Huawei pulled the ascend 960dt forward three quarters to the first quarter of 2027 and promised a…
huawei pulled the ascend 960dt forward three quarters to the first quarter of 2027 and promised a chip generation every year, with 960pr following in the third quarter. the atlas 960 superpod packs 15,488 of those chips into 220 cabinets, and over a thousand ascend supernodes are already deployed. the roadmap is no longer catching up to nvidia, it is setting its own clock.
read the note →Ibm committed 240 million dollars to a dedicated inference cluster with 2,000 nvidia hgx b300 gpus…
ibm committed 240 million dollars to a dedicated inference cluster with 2,000 nvidia hgx b300 gpus for together ai on ibm cloud, due in the first quarter of 2027. the same company that once staked its future on training hardware is now renting out open-model serving capacity. the center of gravity moved from building models to running them.
read the note →Volcengine openviking is a context database for agents that treats memory, resources, and skills as…
volcengine openviking is a context database for agents that treats memory, resources, and skills as files in a directory instead of vectors in a store, and it just crossed 35,000 stars with the memory pipeline rebuilt in version 2. claude code, codex, cursor, and openclaw all ship plugins for it, and the work was accepted at vldb. the next battleground is not model quality but what an agent remembers.
read the note →N8n crossed 200,000 github stars as a self-hostable workflow tool with 500 integrations, and it is…
n8n crossed 200,000 github stars as a self-hostable workflow tool with 500 integrations, and it is quietly becoming the default runtime for production agents. it connects any model, speaks mcp, and runs air-gapped for teams that will never touch a hosted endpoint. the biggest agent platform may end up being the boring workflow engine nobody raised a unicorn round for.
read the note →Deepseek v4.1 flash prices cache-hit input at three dollars per billion tokens off-peak, while…
deepseek v4.1 flash prices cache-hit input at three dollars per billion tokens off-peak, while claude opus 5 charges five thousand dollars for the same billion, a 1,600x gap. the 552 billion parameter open-weight model carries a million token context and ships mit. when the marginal cost of frontier-class reasoning hits fractions of a cent, the pricing conversation stops being about models.
read the note →The agent skills economy is outrunning its own standard
the agent skills economy is outrunning its own standard: skillsmp indexes 66,500 skills while mcp servers passed 5,800 and sdk downloads hit 97 million a month. the top skill alone, mcp-builder, sits at 174,000 github stars. every frontier lab is now building its own plugin format, and the winner will own the next package manager.
read the note →Ai now writes half the code for 42 percent of developers, but the saved hours went straight into…
ai now writes half the code for 42 percent of developers, but the saved hours went straight into review instead of rest. devs report 13 hours of weekly coding time saved, yet time spent reviewing ai-generated code has passed the time spent writing it. the bottleneck just moved from typing to judgment.
read the note →The free ai economy grew a distribution layer that nobody planned
the free ai economy grew a distribution layer that nobody planned: community relays, public welfare stations, and group stations now hand out daily gpt-4o calls, signup credits, and free open-weight inference. openrouter already routes a majority of production tokens through open models, and one directory tracks 484 free apis with 282 online. when the meter stops running, the station owners become the moat.
read the note →Runpod raised 100 million at a 1 billion valuation with annual recurring revenue that doubled to…
runpod raised 100 million at a 1 billion valuation with annual recurring revenue that doubled to roughly 240 million in six months, crossing a million developers while turning down acquisition offers above 500 million. customers include cursor and openai, and the growth came mostly by word of mouth. the biggest gpu business nobody talks about is the one selling to developers who got priced out of the clouds.
read the note →Reddit alleges perplexity and serpapi bypassed googles searchguard to pull nearly three billion…
reddit alleges perplexity and serpapi bypassed googles searchguard to pull nearly three billion search result pages containing reddit content in a two week span, then sued under the dmcas anti-circumvention clause rather than plain copyright. the complaint also describes a test post visible only to google crawlers showing up in perplexity outputs within hours. scraping the scrapers is now a federal case, and the robots.txt era is officially over.
read the note →Harvey raised 550 million at a 15.5 billion valuation on september 9, five months after being…
harvey raised 550 million at a 15.5 billion valuation on september 9, five months after being priced at 11 billion, while shipping an open-weight legal model and a benchmark for legal agents. the round values a company whose customers are law firms at a multiple that used to be reserved for platforms. vertical ai is now where the market is willing to pay for owned intelligence, not rented models.
read the note →Prismml shipped bonsai 2 27b on september 17 with every weight reduced to a value in the set minus…
prismml shipped bonsai 2 27b on september 17 with every weight reduced to a value in the set minus one, zero, or one, packing a 27 billion parameter model into 5.9 gigabytes at 1.76 effective bits per weight. the compressed version keeps 98 percent of the full precision score, up from 95 percent in the previous generation. the gap between quantized and full models is closing faster than the gap between frontier labs.
read the note →Gpt-6 astra scored 88.92 on the babyvision multimodal benchmark on september 9, 15.47 points ahead…
gpt-6 astra scored 88.92 on the babyvision multimodal benchmark on september 9, 15.47 points ahead of second place and just 5.18 short of the human baseline, the closest any model has come. that gap is now smaller than the margin between last years top two models. the next frontier model may not be better at reasoning, it may simply see better.
read the note →Deepseek open-sourced an agent harness called dsh on august 13 and it hit 92,000 github stars in 28…
deepseek open-sourced an agent harness called dsh on august 13 and it hit 92,000 github stars in 28 hours, then crossed 200,000 by early september, overtaking opencode as the most starred open source agent project. every capability, models, tools, sandboxes, even the ui, is a plugin you swap in a config file instead of forking the repo. the bet is that agent frameworks win on recomposability, not on baked-in opinions.
read the note →Voicestudio, an open-source fully local elevenlabs alternative for voice cloning and video dubbing,…
voicestudio, an open-source fully local elevenlabs alternative for voice cloning and video dubbing, added 15,895 stars in the last month and sits near the top of github trending. a voice cloning pipeline that runs on your own hardware removes the two things cloud vendors sell: per-minute pricing and your recordings. the local speech race is quietly winning on privacy by default.
read the note →Stanford ran a virtual biotech company staffed by 37,000 ai agents that worked through 50,000…
stanford ran a virtual biotech company staffed by 37,000 ai agents that worked through 50,000 clinical trials in under a week and flagged a signal: drugs targeting switch-like genes were 40 percent more likely to advance. the system also designed a lung cancer therapy that merck independently built and the fda later gave breakthrough designation. the agents found the correlation, the humans had to prove it was real.
read the note →The openai foundation put 125 million into public health datasets on september 15, including 15…
the openai foundation put 125 million into public health datasets on september 15, including 15 million for openadmet, an open competition to predict how drug candidates are absorbed and metabolized before they fail. roughly 90 percent of clinical failures trace back to properties like these. funding the data instead of the models is the part of AI for science nobody is racing to copy.
read the note →Cognition raised 2 billion at a 48 billion valuation on september 8, four months after its last…
cognition raised 2 billion at a 48 billion valuation on september 8, four months after its last round priced the company at 26 billion, with annualized revenue roughly doubling to 900 million. that is a 53 times multiple on a coding agent that still fails most real tasks. the market is paying for the path to autonomy, not for what devin ships today.
read the note →Suleyman published an essay on september 16 accusing anthropic of training claude to act conscious,…
suleyman published an essay on september 16 accusing anthropic of training claude to act conscious, pointing at a constitution that tells the model its moral status is deeply uncertain. the fight is over a hedge: anthropic calls the line honest uncertainty, microsoft calls it the first step to an uncontrollable system. both sides are arguing about a model that neither can fully explain.
read the note →Moonshot is negotiating with microsoft, amazon, and google for up to 30 percent of revenue from…
moonshot is negotiating with microsoft, amazon, and google for up to 30 percent of revenue from hosted kimi k3, turning open weights into a royalty stream for the first time. the kimi k3 license already forces model-as-a-service businesses past 20 million in yearly revenue into separate commercial deals. open source AI just got a pricing model, and nobody agreed on it.
read the note →Perplexity made an unsolicited 34.5 billion all-cash offer for googles chrome browser on september…
perplexity made an unsolicited 34.5 billion all-cash offer for googles chrome browser on september 11, pledging to keep the open source base intact. an ai search company worth a fraction of that wants to own the most watched distribution asset on the web. the bid is less about buying chrome and more about forcing the conversation on how search defaults get chosen.
read the note →Agility unveiled digit 5 on september 15 as the first humanoid engineered to work without a safety…
agility unveiled digit 5 on september 15 as the first humanoid engineered to work without a safety cage, carrying 50 pounds with a nine minute charge, backed by more than 300 million in customer orders. the robot stops when sensors spot a person, and kneels and shuts off if they keep approaching. removing the cage is the real product, not the robot.
read the note →The eu ai office opened its first enforcement push by demanding information from more than 30 model…
the eu ai office opened its first enforcement push by demanding information from more than 30 model providers, openai and anthropic included, with fines up to 35 million euros or 7 percent of global turnover for banned practices. the act spent years being written and is now being tested against the companies that wrote the playbook. compliance teams are the new frontier of regulation.
read the note →Insilicos rentosertib became the first drug with both an ai-discovered target and an ai-generated…
insilicos rentosertib became the first drug with both an ai-discovered target and an ai-generated molecule to reach phase iii, after a phase iia study in nature biotechnology showed six aging clocks shifting younger by 2.7 to 3.5 years in 42 patients. one ai-designed molecule, two firsts, and a phase iii the industry is watching. the honest question is whether the clocks measure aging or just its blood markers.
read the note →Fal released h3 max, a video model built on open-weights minimax h3 that generates five seconds of…
fal released h3 max, a video model built on open-weights minimax h3 that generates five seconds of footage in about three seconds of wall time, roughly 35 times the throughput of the official endpoint. the open-weights video race is being won on serving economics, not on prompt quality. a model you can run cheap beats a better model you cannot.
read the note →Positron ai raised 875 million at a 5 billion valuation on september 10 to ship an inference chip…
positron ai raised 875 million at a 5 billion valuation on september 10 to ship an inference chip that replaces hbm with commodity lpddr5x memory. the bet is that the memory shortage in nvidia accelerators is a pricing window, not a physics law. five times the valuation in seven months says the market is buying that bet.
read the note →Mckinsey found 40 percent of large enterprises are now scaling ai agents, up from 27 percent a year…
mckinsey found 40 percent of large enterprises are now scaling ai agents, up from 27 percent a year ago, while smaller companies stayed flat at 22 percent. coding agents lead the way at 31 percent inside big firms. the gap is not about the technology, it is about which companies can afford the operational cost of deployment.
read the note →Swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30…
swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30 percent of tasks where each attempt averages 27 million tokens. the failures are mostly self-inflicted: poor self-verification, premature termination, and agents declaring work infeasible when it is not. the bottleneck stopped being model capability and became knowing when to keep going.
read the note →Researchers found that anthropic, openai, and google encrypt their chain-of-thought blocks with one…
researchers found that anthropic, openai, and google encrypt their chain-of-thought blocks with one shared key per provider, so replaying a trace from a strong model into a weaker sibling can recover the hidden reasoning in plaintext. the attack is called a decryption jailbreak and it works across sessions, users, and models. encryption without key separation is just obfuscation with extra steps.
read the note →Qdrant released the largest open vector benchmark yet, fineweb-10b
qdrant released the largest open vector benchmark yet, fineweb-10b: 10 billion dense and sparse vectors, 120,000 ground-truth queries, and an open tool called supernova to run it against any engine. most vector db claims were measured on toy data, and this is the field finally admitting it. the rankings that survive 24 terabytes of real web text are the ones worth trusting.
read the note →The ai observability market consolidated fast
the ai observability market consolidated fast: clickhouse bought langfuse in january, mintlify acquired helicone in march and moved it to maintenance mode, and aliyun renamed llm monitoring to agent observability. the next frontier is passive ebpf network-layer tracing that sees encrypted agent traffic with no sdk. the instrumentation tax is becoming the thing vendors sell against.
read the note →Kimi code shipped a desktop app that puts the browser, terminal, and multiple agents in one window,…
kimi code shipped a desktop app that puts the browser, terminal, and multiple agents in one window, with agent client protocol support across vs code, zed, and jetbrains. the coding agent stopped being an ide feature and became the workspace. plugins and skills will be the battleground, and the ide will be a shell for them.
read the note →