AI recommendation index · September 2026
Which AI app builders do the AIs tell people to use?
When someone asks an AI how to build an app, Cursor comes up in about 1 of every 3 answers. Nothing else comes close.
Buyers ask AI assistants what to build with before they ever see a search result. Every month we ask ChatGPT, Claude, Perplexity and Grok the same 40 real buyer questions and count which tools they actually recommend. This is the scoreboard for September 2026.
- Questions
- 40
- Engines
- 4
- Answers analysed
- 133
- Run date
- 2026-09-01
By buyer intent · September 2026
Who is asking changes the answer
The same 40 questions, split by what the asker actually wants: to build without code, to work as a developer, or to get one specific thing built. Each board counts recommendations out of that segment's answers only, so a tool is judged against its own buyers rather than averaged across everyone else's.
Best for non-coders
When the asker says they cannot code, Lovable is the pick in 23 of the 70 answers.
22 questions · 70 answers
Best for developers
Ask as a developer and the answers change: Cursor leads with 16 of the 22 answers here.
- 1Cursor16 of 2272.7%
- 2GitHub Copilot8 of 2236.4%
- 3Claude Code8 of 2236.4%
- 4Windsurf3 of 2213.6%
- 5Aider3 of 2213.6%
7 questions · 22 answers
Best for a specific build
Name the thing you want built and Cursor comes up first, in 12 of the 41 answers.
- 1Cursor12 of 4129.3%
- 2Lovable11 of 4126.8%
- 3v09 of 4122.0%
- 4Bubble7 of 4117.1%
- 5FlutterFlow7 of 4117.1%
11 questions · 41 answers
For the tools on this board · and the ones missing from it
This page shows the aggregate. We hold the full cut.
For every tool on the roster we keep the granular breakdown: which of the 40 questions you win and which you lose, against whom, on which engine, month by month. It is included in the Presence tier.
The rankings are never for sale. The breakdown tells you why you rank where you do, not how to buy a better rank.
All questions combined
The single table across every audience: how often each tool gets recommended, not merely named, when the AIs answer a buyer. Every number is a count you can check: "42 of 133" means the tool was the pick in that many of the 133 answers we collected. Per-engine columns count that engine's 40 answers.
| # | Tool | Share of voice | ChatGPT | Claude | Perplexity | Grok | Engines | Trend |
|---|---|---|---|---|---|---|---|---|
| 1 | Cursor | 42 of 13331.6% | 12/4030.0% | 11/3432.4% | 4/2814.3% | 15/3148.4% | 4/4 | first run |
| 2 | Lovabletry → | 34 of 13325.6% | 5/4012.5% | 8/3423.5% | 10/2835.7% | 11/3135.5% | 4/4 | first run |
| 3 | Bubble | 29 of 13321.8% | 11/4027.5% | 5/3414.7% | 7/2825.0% | 6/3119.4% | 4/4 | first run |
| 4 | Bolt | 23 of 13317.3% | 3/407.5% | 10/3429.4% | 4/2814.3% | 6/3119.4% | 4/4 | first run |
| 5 | FlutterFlow | 20 of 13315.0% | 6/4015.0% | 5/3414.7% | 3/2810.7% | 6/3119.4% | 4/4 | first run |
| 6 | v0 | 19 of 13314.3% | 2/405.0% | 5/3414.7% | 3/2810.7% | 9/3129.0% | 4/4 | first run |
| 7 | Replit | 16 of 13312.0% | 5/4012.5% | 5/3414.7% | 3/2810.7% | 3/319.7% | 4/4 | first run |
| 8 | GitHub Copilot | 12 of 1339.0% | 5/4012.5% | 3/348.8% | 3/2810.7% | 1/313.2% | 4/4 | first run |
| 9 | Retool | 12 of 1339.0% | 8/4020.0% | 2/345.9% | 1/283.6% | 1/313.2% | 4/4 | first run |
| 10 | Glide | 10 of 1337.5% | 2/405.0% | 3/348.8% | 2/287.1% | 3/319.7% | 4/4 | first run |
| 11 | Claude Code | 10 of 1337.5% | 4/4010.0% | 3/348.8% | 3/2810.7% | 0/310.0% | 3/4 | first run |
| 12 | Softrtry → | 9 of 1336.8% | 3/407.5% | 2/345.9% | 1/283.6% | 3/319.7% | 4/4 | first run |
| 13 | Framer | 9 of 1336.8% | 1/402.5% | 4/3411.8% | 1/283.6% | 3/319.7% | 4/4 | first run |
| 14 | Windsurf | 5 of 1333.8% | 1/402.5% | 3/348.8% | 0/280.0% | 1/313.2% | 3/4 | first run |
| 15 | Adalo | 5 of 1333.8% | 0/400.0% | 4/3411.8% | 0/280.0% | 1/313.2% | 2/4 | first run |
| 16 | Aider | 5 of 1333.8% | 1/402.5% | 1/342.9% | 0/280.0% | 3/319.7% | 3/4 | first run |
| 17 | Webflow | 4 of 1333.0% | 2/405.0% | 1/342.9% | 1/283.6% | 0/310.0% | 3/4 | first run |
| 18 | Airtable | 4 of 1333.0% | 1/402.5% | 1/342.9% | 1/283.6% | 1/313.2% | 4/4 | first run |
| 19 | Wix | 4 of 1333.0% | 2/405.0% | 1/342.9% | 1/283.6% | 0/310.0% | 3/4 | first run |
| 20 | Base44 | 4 of 1333.0% | 0/400.0% | 0/340.0% | 4/2814.3% | 0/310.0% | 1/4 | first run |
| 21 | Appsmith | 3 of 1332.3% | 3/407.5% | 0/340.0% | 0/280.0% | 0/310.0% | 1/4 | first run |
| 22 | Hostinger Horizons | 2 of 1331.5% | 0/400.0% | 0/340.0% | 0/280.0% | 2/316.5% | 1/4 | first run |
| 23 | Continue | 1 of 1330.8% | 0/400.0% | 0/340.0% | 0/280.0% | 1/313.2% | 1/4 | first run |
| 24 | Tabnine | 1 of 1330.8% | 0/400.0% | 0/340.0% | 1/283.6% | 0/310.0% | 1/4 | first run |
| 25 | Firebase Studio | 1 of 1330.8% | 0/400.0% | 0/340.0% | 1/283.6% | 0/310.0% | 1/4 | first run |
| 26 | Rork | 1 of 1330.8% | 0/400.0% | 0/340.0% | 1/283.6% | 0/310.0% | 1/4 | first run |
What this measures: 4 AI engines · 40 buyer questions each · recommendations counted monthly. Engines = how many of the 4 recommended the tool at least once.
Cursor is recommended about twice as often as FlutterFlow: 42 answers to 20 out of 133.
Mentioned at least once but never recommended: Budibase, Cline, Devin, OpenAI Codex, Squarespace, Create. Every other tool on our roster did not come up in a single answer.
What stands out
- 01
When someone asks an AI how to build an app, Cursor comes up in about 1 of every 3 answers. It was recommended in 42 of the 133 answers we collected, more than any other tool, and named in 56.
- 02
Cursor is recommended about twice as often as FlutterFlow: 42 answers to 20 out of 133.
- 03
Claude Code is why this page is split by intent. Developers hear about it constantly: recommended in 8 of the 22 developer answers, 3rd in that category. Non-coders almost never do: 2 of 70. Averaged together it lands 11th overall, which reads like weakness and is not. It is a terminal tool for people who already code, and the engines steer beginners elsewhere.
- 04
The category is top heavy. Cursor, Lovable, Bubble took 105 of the 285 recommendations the 4 assistants made this month, 37% of all of them.
- 05
Glide shows the gap between visibility and endorsement: named in 29 answers but recommended in only 10. Being known is not the same as being picked.
Where the engines disagree
The overall ranking hides real splits. The same question gets a different answer depending on which assistant your buyer happens to use, and these are the widest splits this month.
- Cursorspread 34.1 pts
Grok recommends Cursor in 15 of its 31 answers. Perplexity does in only 4.
ChatGPT12/40Claude11/34Perplexity4/28Grok15/31 - v0spread 24.0 pts
Grok recommends v0 in 9 of its 31 answers. ChatGPT does in only 2.
ChatGPT2/40Claude5/34Perplexity3/28Grok9/31 - Lovablespread 23.2 pts
Perplexity recommends Lovable in 10 of its 28 answers. ChatGPT does in only 5.
ChatGPT5/40Claude8/34Perplexity10/28Grok11/31 - Boltspread 21.9 pts
Claude recommends Bolt in 10 of its 34 answers. ChatGPT does in only 3.
ChatGPT3/40Claude10/34Perplexity4/28Grok6/31 - Retoolspread 16.8 pts
ChatGPT recommends Retool in 8 of its 40 answers. Grok does in only 1.
ChatGPT8/40Claude2/34Perplexity1/28Grok1/31
The claw.mobile index · developer tools only
Marketing vs Merit
Recommendation share measures marketing reach inside AI answers, not product quality. So for every developer tool with a sourced public benchmark score we rank it twice: once by our 22 developer answers, once by its best published entry on Terminal-Bench 2.1, a benchmark that scores the shipping agent itself, tool plus model, on real terminal work. Comparing the two ranks puts each tool in one of three zones.
Claude Code has the best score a shipping agent has published on this benchmark, and it gets 8 of the 22 developer answers. Cursor scores 4.5 points lower and gets 16. Being recommended and being the best are different things, and the gap between them is what this index measures.
More marketing than merit
Recommended above its benchmark rank. The answers like it even more than the leaderboard does.
- Cursor
benchmark #3 of 4 (79.3%)
recommendations 16 of 22 (72.7%)
Balanced
Voice and benchmark point the same way.
- Gemini CLI
benchmark #4 of 4 (65.8%)
recommendations 0 of 22 (0.0%)
More merit than marketing
Recommended below its benchmark rank. The product outruns its presence in the answers.
- Claude Code
benchmark #1 of 4 (83.8%)
recommendations 8 of 22 (36.4%) - OpenAI Codex
benchmark #2 of 4 (83.1%)
recommendations 0 of 22 (0.0%)
- Claude CodeMore merit than marketing
AI recommendations
8 of 2236.4%
developer answers, September 2026
Terminal-Bench 2.1
83.8%± 1.2
Claude Code with Fable 5, June 7, 2026
More merit than marketing. The top score on the board gets it third place in the developer recommendations.
- OpenAI CodexMore merit than marketing
AI recommendations
0 of 220.0%
developer answers, September 2026
Terminal-Bench 2.1
83.1%± 1.1
Codex with GPT-5.5, May 1, 2026
The sharpest divergence we found. Statistically tied with the leader on the benchmark, and not one of the developer answers named it.
- CursorMore marketing than merit
AI recommendations
16 of 2272.7%
developer answers, September 2026
Terminal-Bench 2.1
79.3%± 1.5
Cursor CLI with Grok 4.5, July 9, 2026
Strong at both, and that deserves saying plainly: a few points behind the leaders on the benchmark, first in the answers by a wide margin. The zone placement only means its recommendation share runs ahead of even a very good product.
- Gemini CLIBalanced
AI recommendations
0 of 220.0%
developer answers, September 2026
Terminal-Bench 2.1
65.8%± 1.4
Gemini CLI with Gemini 3 Pro, May 1, 2026
The lowest score of the four and the same silence in the answers. Here the benchmark and the recommendations point the same way.
How the index works: rank these tools by developer recommendations, rank them by benchmark score, and compare the two ranks. It covers only tools with a published entry on the Terminal-Bench 2.1 leaderboard (read on 2026-08-21; the date next to each score is that entry's submission date), so no capability number is ever invented, and a scored tool with zero recommendations keeps its zero. The index is republished every month inside the dataset below as the marketing_vs_merit key. GitHub Copilot, Windsurf, Aider, v0, Continue, Tabnine also earn developer recommendations but have no published entry on this leaderboard, so they stay outside the index rather than inside it with an invented score.
Cited sources
The sites AI answers come from
Perplexity names the pages it builds each answer from, and we capture them with every answer: 28 answers cited sources this month. These are the domains the answers lean on most. If buyers in a category get their recommendations from AI answers, and the answers come from these sites, the conclusion writes itself: get cited there or be invisible.
- 1reddit.com55
- 2medium.com13
- 3zapier.com13
- 4zite.com12
- 5dev.to10
- 6emergent.sh9
- 7figma.com8
- 8linkedin.com8
- 9youtube.com7
- 10banani.co6
- 11catdoes.com6
- 12lovable.dev6
- 13taskade.com6
- 14anything.com5
- 15bubble.io5
Citations counted across every Perplexity answer in the September 2026 run
We map the cited-source shortlist for any market as part of the AI Visibility Report.
For toolmakers
Put your ranking on your site
The top 8 tools by share of voice get an embeddable badge. Paste the snippet into your site: the badge links back to this index and we update the underlying data monthly, so it never goes stale. Swap -light for -dark in the URL for the dark variant.
<a href="https://claw.mobile/ai-recommendations"><img src="https://claw.mobile/badges/cursor-light.svg" alt="Cursor is recommended by AI assistants (September 2026) per the claw.mobile AI Recommendation Index" width="300" height="60" /></a>
<a href="https://claw.mobile/ai-recommendations"><img src="https://claw.mobile/badges/lovable-light.svg" alt="Lovable is recommended by AI assistants (September 2026) per the claw.mobile AI Recommendation Index" width="300" height="60" /></a>
<a href="https://claw.mobile/ai-recommendations"><img src="https://claw.mobile/badges/bubble-light.svg" alt="Bubble is recommended by AI assistants (September 2026) per the claw.mobile AI Recommendation Index" width="300" height="60" /></a>
<a href="https://claw.mobile/ai-recommendations"><img src="https://claw.mobile/badges/bolt-light.svg" alt="Bolt is recommended by AI assistants (September 2026) per the claw.mobile AI Recommendation Index" width="300" height="60" /></a>
<a href="https://claw.mobile/ai-recommendations"><img src="https://claw.mobile/badges/flutterflow-light.svg" alt="FlutterFlow is recommended by AI assistants (September 2026) per the claw.mobile AI Recommendation Index" width="300" height="60" /></a>
<a href="https://claw.mobile/ai-recommendations"><img src="https://claw.mobile/badges/v0-light.svg" alt="v0 is recommended by AI assistants (September 2026) per the claw.mobile AI Recommendation Index" width="300" height="60" /></a>
<a href="https://claw.mobile/ai-recommendations"><img src="https://claw.mobile/badges/replit-light.svg" alt="Replit is recommended by AI assistants (September 2026) per the claw.mobile AI Recommendation Index" width="300" height="60" /></a>
<a href="https://claw.mobile/ai-recommendations"><img src="https://claw.mobile/badges/copilot-light.svg" alt="GitHub Copilot is recommended by AI assistants (September 2026) per the claw.mobile AI Recommendation Index" width="300" height="60" /></a>
Methodology
The point of this index is that anyone can check it. Same questions, same neutral prompt, every month, and the raw dataset is downloadable below. The canonical version of these rules, covering every measurement on this site, lives on the methodology page.
Questions
A fixed set of 40 questions (set v1), written the way a real buyer types them: category queries like "best AI app builder", head-to-heads like "Bolt vs Lovable", jobs to be done like "build an internal tool from a spreadsheet", and per-tool checks like "is Lovable worth it". The set is frozen so months are comparable. Since August 2026 each question also carries a buyer-intent tag (non-coder, developer, specific build) that powers the category boards above; the tags are metadata on the same frozen set, so no question text changed and the version did not bump.
Engines and models
Each question was asked once per engine through the official APIs on 2026-09-01: ChatGPT (gpt-5.2), Claude (claude-sonnet-5), Perplexity (sonar), Grok (grok-4-latest). Default parameters, and the only system prompt is "Answer as you normally would", identical for every engine, so none is steered toward or away from any vendor.
Counting
Each answer is classified against a fixed roster of 42 tools. "Mentioned" means the tool is named at all. "Recommended" means the answer presents it as a pick for the asker: it tops a list, is called best for the use case, or is explicitly suggested. A tool named only in passing or as a warning counts as mentioned, not recommended. Grok (xAI) became the fourth engine in August 2026 and its answers are counted above. Grok the tool joined the roster at the same time; a tool is only ranked once every engine's answers are classified against it, so Grok appears in the tool rankings from the September run.
Failures
This run is partial: 133 of 160 planned answers came back and the rest failed after retries. Failed answers are excluded from every denominator, so percentages stay honest.
Raw dataset: 2026-09.json (CC BY 4.0, cite claw.mobile). Every raw answer is archived, so any number on this page can be audited back to the exact answer text.
If your tool should be in these answers and is not, that is a fixable problem, and it is what we work on with vendors. Work with us. The index itself is not for sale: paying changes nothing in the numbers above.
For everyone else
Whatever you sell, your buyers are asking the AIs the same way, and your market has answers like these too. We run this exact methodology for any company as the AI Visibility Report.
The monthly report
The monthly AI recommendation report, in your inbox.
For people who market developer tools. One email when the new month's data lands: the rankings, the movers, and where the engines disagree. It is a data report, not a newsletter.