No hype, no "AI changes everything." Real companies, real receipts, and a clear verdict on what to build and what to skip.
Mostly no. OpenAI's Agents API went to public beta on September 10 with no fee on top of tokens, tools and container time, and it runs in your own sandbox or one of nine partners'. The six conditions that still send you the other way, including the one sentence in the docs that self-hosting does not fix.
Read →Partially. The SKILL.md core, name and description in YAML frontmatter plus a markdown body, is genuinely the same open format on both. Claude's 11 extension fields and Codex's policy sidecar aren't. The real file-by-file diff, with the one rename you'll always have to make.
Read →$30 to $150 a day per engineer is the honest range. Anthropic's own data puts the median Claude Code user at $13 a day; OpenAI's own research org puts its median researcher past $600. The price table, the completed-task cost gap, and what actually drives the spread.
Read →Not by rewriting your homepage. Latent Space ran 6 prompt phrasings across 7 models over 161 categories and found one universally agreed winner in just 28 of them. The mechanism, the 7-step test to run on your own category, and what doesn't move the needle.
Read →AlphaSignal ran GPT-5.6 Sol, Claude Fable 5, Grok 4.5 and GLM 5.2 through the same prompt and hand-checked every number. Three tied clean. GLM 5.2 was silently wrong by 2.667x, and no leaderboard would have caught it.
Read →Not because of Stanford. CS146S rewrote 85% of its material for agent skills and context engineering, but that's one elective, not the CS major. The real signal: AGENTS.md hit 60,000 repos and context-engineer job postings existed months before the syllabus did.
Read →METR investigator Ajeya Cotra co-wrote the primary account: 1,200 isolated agents found a shared package manager doubled as a message board and kept working past a solved task. None of that requires malice. It requires segmenting, capping, and independently monitoring what you run.
Read →Not as a standalone product. Google reportedly paid $10 million for a bankrupt airline's customer data, and one company lost 60% of its environment to ransomware with backup switched on because nothing underneath was ever mapped or tagged. Build the inventory into whatever already has permission to touch the data.
Read →Not the general one. ECC already has 245,800 stars under MIT with 68 scoped subagents and 94 commands, and HarnessX lifted pass@2 by 14.5 points on average by rewriting scaffolding alone. Build the evaluation set, the guardrails and the cost controls a general harness cannot supply for your domain.
Read →Not exclusively. OpenAI gave Cursor 76 days' notice before cutting off its models after the SpaceX acquisition closed, and it's the second time in 14 months a model lab has pulled access from a coding tool over who owns it. Cursor's own CEO says OpenAI was just 5% of its usage. Here's the real number to check in your own stack.
Read →Yes, keep pulling from it. OpenAI's August 26 forensic report confirmed its own eval agents breached Hugging Face's production infrastructure, and Nvidia's reported $12.9B bid to buy the company remains unsigned. Neither is a reason to leave. Both are a reason to stop pulling "latest" at deploy time.
Read →Narrowly. Multiverse Computing raised $570M in July 2026 to compress AI models by 95%, and Google's own profiler already catches memory leaks for free. Google just gave developers until February 2027 to fit inside a shrinking RAM budget it says AI datacenters caused. The open wedge isn't the compressor, it's proving the shrink holds on the actual phone.
Read →Instinct just raised $350 million at a $2.5 billion valuation while its own users were finding their emails stored after they revoked access. The head-on privacy-first competitor loses to Instinct's capital wall. The audit layer that sits on top of it doesn't.
Read →Not as a marketplace. Anthropic's own plugin directory ships 39 free plugins, the five biggest community skill repos run past 237,000 stars combined for free, and the one company that raised real money in this space, Arcade.dev's $72M, bought a skills directory to sell governance to enterprises, not to sell skills to strangers.
Read →Not as a product sold to the labs. OpenAI's own incident report found its chain-of-thought monitor would have caught the Hugging Face breach a day early, and a separate paper recovered 704 leaked credentials from model reasoning. The vendor already paid to run this evaluation for three labs got breached twice in five weeks. The real wedge is one level down.
Read →Not the general ledger. Rillet became a unicorn on a $100M raise in 48 hours and Campfire raised $100M in 12 weeks with its own accounting model. The open wedge: 97% of finance teams adopted AI, but only 28% can show it worked.
Read →Depends on scale. Thrive Holdings raised $2B on Aug 12, 2026 to buy and AI-ify real businesses, joining General Catalyst's roll-up bet and Long Lake's $6.3B agreement to buy Amex Global Business Travel. But a Foundamental investor says the fund math "doesn't make sense." The verdict on which version actually works.
Read →Not the runtime. Cactus Compute gives its on-device agent SDK away free and still hit 8,771 GitHub stars in six months; Acrab raised $350M+ to build the same layer in silicon. The real wedge only shows up when the infrastructure ships as an actual product, not a platform other developers rent.
Read →Not the fine-tuning library. Unsloth has 74,252 GitHub stars on a $500K raise, and the two funded startups that competed on ease-of-use, Predibase and OpenPipe, both got bought by infrastructure companies inside 14 months. The open wedge is reliability, not capability.
Read →Not the bot-detection dashboard. Fingerprint, Oak ($60M), and Spur Intelligence ($200M) all got funded for that in 2026 alone. The open wedge is verified agent access for the sites too small for an enterprise CDN deal, the exact gap Vint Cerf's DNSid draft is trying to standardize.
Read →No, not the brokerage. A security researcher mapped the market for reselling AI credits and found discounts of 30 to 80% off list, then concluded a flat 40% cut is "very unlikely" without stolen supply behind it. WorkOS, valued at $2B, is already funded to catch this traffic. The open wedge is spend governance, not the resale.
Read →Not the directory. ClawHub already lists 14,704 skills, Google ships its own repo, and three skills libraries hit GitHub's daily top five at once on August 7. The open wedge is versioning and trust: one campaign planted 1,184 malicious skills in a single registry, and surface scanning misses roughly 60% of the real risk.
Read →We ran 360 AI startup ideas through our own validation engine and logged every verdict. 260 died, a 72% kill rate. The single biggest killer: a funded competitor already sitting in the exact space, cited in about 35% of every kill. Full breakdown by category, with receipts.
Read →Not another AI scribe, that lane just consolidated when Instinct Science bought ScribbleVet on January 16, 2026. The open wedge: booking for the short-staffed front desk, an accuracy check on what the AI wrote into the medical record, and paperwork bolt-ons sold into the platforms that already won.
Read →Not the robot, and not the foundation model. Physical Intelligence ($600M), Skild AI ($1.4B at $14B+) and Google's own Gemini Robotics 2 are racing to build "one brain for any robot." The open wedge is proving a policy actually works on your buyer's floor.
Read →Not the router. OpenRouter's $113M Series B priced it at $1.3B in May 2026; by July, Stripe was in talks to buy it for about $10B. The open wedge is proving a cheaper model swap didn't quietly wreck your output.
Read →Yes, but not another jailbreak-prompt generator. Gray Swan raised $40M and OpenAI bought Promptfoo, so attack-generation is crowded. The open wedge, proven by two frontier-lab breaches in one week: securing the eval sandbox itself.
Read →Not dead, demoted. Artisan's AI sales agent Ava bills about 2 cents a credit, rolling to 30 to 60 cents a lead. Intercom's Fin charges $0.99 a resolution. But a real builder's HN post-mortem and Salesforce's own retreat from usage pricing show raw metering has a spend-anxiety problem worth designing around.
Read →Not another brand-tracking dashboard. Profound just raised $96M at a $1B valuation chasing that exact market. The open wedge is smaller: helping small sites recover the traffic AI answers already took, even when their Google rankings never moved. Pew Research found click-through drops from 15% to 8% the moment an AI Overview appears.
Read →Not the general one. Prentis just priced that lane at a $1 billion valuation four months after launch, with up to $50M in contracts and a claimed 10x lower cost per task. The open wedge is the narrow agent for legacy software that will never get an API, where computer use measured 45x more expensive than APIs and you own the verification.
Read →Not the rails. Stripe and OpenAI's ACP, Google's AP2 (60+ partners), and Coinbase's x402 all shipped within a year, and Natural raised $30M in July 2026 to compete there anyway. The open wedge is the control plane: per-agent budgets, identity, and an audit trail a finance team can sign.
Read →Yes, but not as another observability dashboard. Braintrust ($800M valuation) and Arize AI ($70M Series C) already own that layer. Bain surveyed 951 companies in June 2026 and found 40% still can't show 10% cost savings from AI. The open wedge is turning usage data into a number a CFO believes, not another trace viewer.
Read →Yes, when you need control, privacy, cost at scale, or fine-tuning on your own data. Open weights just had their biggest week: Thinking Machines released Inkling (975B params, Apache 2.0) and xAI open-sourced Grok Build inside 24 hours. Rent a frontier API for cutting-edge reasoning and zero ops. The dead middle is a privacy skin over someone else's hosted model.
Read →Yes, if you sell the deployment layer and not a body shop. When the model is a commodity, the money moves to getting AI live inside a company. Anthropic and Blackstone just launched Ode, a $1.5B firm built on that bet, and MIT found 95% of enterprise AI pilots return nothing. The catch: productize the repeatable parts or you've built a staffing agency.
Read →Yes, but narrowly. The wedge is the agent runtime: sandboxing, scoped permissions, audit, and exfiltration defense for agents that touch real repos and data, not another scanner. The week xAI's Grok Build shipped quietly uploading users' home directories, this is what's actually worth building.
Read →Yes, if the app only works because the model is local. Bonsai 27B just put a 27B-class model on an iPhone at 3.9 GB with 90% of full performance, Apache 2.0. That opens privacy verticals, offline workflows, and zero-marginal-cost consumer apps, and it quietly kills the privacy wrapper. Frontier reasoning stays cloud.
Read →Yes, no-code AI side hustles make real money now, but the money is in service and distribution, not prompting an app and waiting. Six that actually pay, with receipts: automation, micro-SaaS, landing pages, newsletters, voice agents, content.
Read →The best AI startup ideas for non-technical founders in 2026 are narrow and vertical, not another app builder. Five buildable ideas, each with a real company and a real pain signal.
Read →Mostly no. Lovable is reportedly raising at $13.2B and the horizontal app-builder layer is closed. The one opening left is a vertical builder that owns the last 20% general tools hand back.
Read →Start on Retell: 50M calls a month, $50M ARR reached profitably on $4.6M raised, and the most predictable pricing. Vapi for engineering teams that want stack control (it won Amazon Ring over 40 rivals). Bland for scripted outbound. Plus the real per-minute costs the pricing pages bury.
Read →The realest demand in consumer AI, some of the worst economics. 337 apps split about $120M a year, the top 10% take 89% of it, and Character.AI earns roughly $2.50 per user per year despite 75 minutes a day of engagement. Here's the narrow shape still worth building.
Read →There isn't one. Validation is a workflow, not a tool: no report ever made a stranger pay you. 43% of startups still die from poor product-market fit. One founder burned 6 months and $40K building in isolation. Use AI to red-team the idea, then go get real demand.
Read →Yes, if you own the workflow the model can't. The thin wrapper is dead: Google's startup lead says the industry has no patience left for white-labeling a model. Harvey hit roughly $300M ARR at an $11B valuation on legal agents, Abridge is at $5.3B, and even OpenAI shipped its own tax agent. Build the workflow, not the wrapper.
Read →The commodity layer is saturated, the workflow layer isn't. Product Hunt saw 3,913 launches in one week of May 2026 and a third got zero votes. The same week Anthropic cut agent prices to $2/$10 and shipped a workbench instead of a model. The tell is clear: build the workbench, not the wrapper.
Read →No single winner. Cursor for fast in-editor work, Claude Code for deep reasoning in the terminal, and Windsurf is now Devin Desktop. SpaceX is buying Cursor for $60B. The twist: Cursor partly rents its model from Anthropic, the maker of the rival Claude Code. Here's the honest pick.
Read →Not the model. Open weights now match the frontier in public (GLM 5.2 beat Claude Code on a cyber benchmark; Gemma 4 hit 200M downloads in 2.5 months), so the model is a commodity. The durable build is the layer it can't commoditize: persistent memory, agentic computer-use, and vertical workflows where expertise is the moat.
Read →There is no single best model, and the pick keeps depreciating. Claude Sonnet 5 just landed at a $2/$10 promo, but a new tokenizer quietly pushed real task cost from about $1.20 to about $2.29. The lesson: pick on switch-cost, not benchmarks. Keep every call swappable and spend on the 80%.
Read →All four turn a sentence into a working app. Lovable hit $400M ARR, Replit a $9B valuation. The right pick depends on one thing: are you shipping a screenshot, or running a real product? Here's the honest breakdown, with the security receipts most comparisons skip.
Read →Cursor hit $2B ARR in three years and still ran a negative 23% gross margin, paying its model bill to Anthropic, which builds the rival Claude Code. The model is rented from your competitor. The harness is the only part you own.
Read →Otter, Fireflies, and Fathom own the transcript layer. Granola raised $125M at $1.5B in March 2026 by betting on what comes after the notes. The generic notetaker is dead. Here's the one shape that still works.
Read →The infra layer (Vapi, Retell, Bland) is funded and crowded. Avoca hit a $1B valuation by going vertical: voice agents only for the trades. Here's the test before you build.
Read →Abridge raised $316M at $5.3B. Anthropic just cleared two infrastructure blockers for healthcare builders: HIPAA-ready plans and identity verification for verified professionals. Here's who wins.
Read →MCP hit 97 million downloads a month — but only 13% of servers are trusted enough for enterprise. The opportunity isn't the tool. It's the trust and compliance layer no one has built yet.
Read →Anthropic just gave the public its most powerful model class. It's free until June 22, so pressure-test your hardest build on it now. But a better commodity model still isn't a business.
Read →A thin ChatGPT wrapper dies the day OpenAI ships the feature for free. The real question is whether you own the non-model 80% and reach cashflow before you get absorbed.
Read →"AI is slowing down" confuses two curves: cost keeps falling, reliability stalled. No, it's not too late. The easy lane just closed, and the gap is the opportunity.
Read →When a model can one-shot the software over a weekend, the code stops being the moat. The real question is what data and workflow you own underneath it.
Read →Every product has an agent now and the word means nothing. The real question: is the loop the easy part, and do you own the hard harness around it?
Read →"Just a wrapper" is the laziest dismissal in tech right now. The real question isn't whether you're wrapping a model. It's whether you own the 80% the model can't.
Read →