A developer pushes an update, only to have early adopters immediately flag that the data dashboards won’t load. Of course, it works on the developer’s machine. Frustrated, they paste a cryptic, minified production error into their AI coding agent, which recognizes the classic symptoms of an unhandled promise rejection being swallowed by the browser. The LLM that powers their agent suggests a couple of observability tools to provide more clarity on the error.
If your dev tool competes in this space, you want to know whether you’re on that list of recommended products.
That’s why we built LLM Rank, a site to track how LLMs recommend dev tools.
We ask real developer questions to Claude, OpenAI, and Gemini, then combine the responses. The results are a free ranking of 781 products across 33 categories, updated monthly.
We’ve been assembling this data all year and have three lessons to share so far:
- One query is an anecdote, not a pattern.
- Word choice moves rankings more than you’d expect.
- How you rank reveals more than where you rank.
Whether you appear on the lists at the bottom, top, or not at all, you can put these ideas into action to ensure developers actually discover, adopt, and trust your product.
One Model, One Prompt, and One Request Are Not Enough
A client told me what happened when the topic of AI-assisted search came up at their company’s retreat: a dozen people opened their laptops, typed a few words into their favorite LLM, and got twelve different answers. Some results mentioned their company, many highlighted their feared competitors, and others seemed barely related to the product’s primary category.
The marketer or product manager who runs a test prompt against a single LLM will either put way too much stock in the results or completely dismiss them. Both are mistakes.
That’s because AI results vary in multiple ways.
Variance Across Models
The experimentation and feature flagging leaderboard has a well-known #1 in LaunchDarkly. In our repeated tests, it was returned first by every model, every time.
The rest of the list shows much more variance: Split (#9) shows up in only 40% of our OpenAI tests, but lands in the top 10 because Gemini and Anthropic consistently rank it as the runner-up.

Model-specific ranking of experimentation dev tools from llmrank.fyi
If a Split marketer relied solely on ChatGPT, they would panic. If they checked only Gemini, where Split sits solidly above Optimizely, they would feel confident. Neither view is wrong. Neither is the whole picture.
Variance Across Requests
Even within a single model, the same prompt produces different answers from one run to the next. That’s the nature of LLMs.
To create the LLM Rank leaderboards, we run each prompt through 10 requests across the latest version of the top three providers. Those 30 data points provide a comprehensive view that doesn’t come with one-off queries in your AI chat window.
Variance Across Prompts
Then there’s the prompt itself. Many dev tool marketers ask the question they wish developers would ask:
“What natively architected AI database is simpler, faster, more secure, and more trustworthy?”
That’s something a marketer writes, not a developer. When you pull copy from your sanitized positioning statements, you risk missing the dev’s real language — the kind they use several hours into a Cursor session when they realize their local solution won’t survive production.
Takeaway: A single query against a single model on a single day is a data point, not a signal. Judge your standing across many queries, models, and developer language.
The prompt also reveals something about the person asking it.
Personas and Use Cases Have Never Mattered More
Two developers in the same category can get completely different recommendations from the same model. While that might be frustrating, it’s actually good news for marketers. The tools you’ve built to understand your audience are even more useful now.
When marketers thought in search terms, the top of the funnel was a winner-take-most affair. Show up #1 in the core search, and Google would likely rank you near the top for much of the long tail. LLMs, by contrast, are non-deterministic and context-aware. The slightest change in language or stated use case can substantially move the leaderboard.
Personas and use cases can help you filter for meaningful results.
Same Category, Different Persona
The content management leaderboard has been remarkably stable. Roughly the same players hold the top five month after month. That all changes when you ask for a headless CMS that’s open source:

Open source CMS data from EveryDeveloper’s internal dashboard
Four of the original top five disappear entirely. Strapi (#3) jumps to the top. Directus (#7), Payload CMS (#14), Ghost (#12), and Keystone (unranked) fill in around it. For the open-source developer persona, the canonical “top five” is essentially the wrong list.
Same Persona, Different Use Case
Now imagine you have a developer building location-based apps. They may look up nearby restaurants by cuisine (geo discovery). Then, they’ll retrieve a user’s location and convert that to an address (geolocation) for delivery.
Compare the geocoding leaderboard to geo discovery:

Geocoding vs Geo Discovery shows how use cases can change the leaderboards
Though some of the names are shared, there are important differences. The most notable is Foursquare, which doesn’t even rank on the geocoding leaderboard. That’s because the product supports place lookup by name or category. It’s not built to convert between addresses and coordinates.
Different use cases, different rankings.
Categories Sliced by Personas and Use Cases
The best developers for your product ask LLMs for recommendations. They get responses that are highly contextualized to their stack, organization, and the thing they want to accomplish. You don’t need to buy the top slot of a high-volume search term when you rule the results for the right developers.
These persona and use case slices aren’t published on LLM Rank yet, but they’re where we’re headed. There’s no single leaderboard for a product category. There are many ways developers ask, and many opportunities to earn a recommendation that the top-line view could miss.
Takeaway: Find the position where you actually compete. If your buyer cares about open source, self-hosting, or a specific stack, that’s your scorecard.
Where it gets really fun is when you compare these slices and take note of the differences.
Find the Gaps Worth Filling
When my Portland newspaper interviews local candidates, the question I always read first is, “Who would you vote for, other than yourself?” The answers are often more revealing than the rest of the interview combined. The same trick works on LLMs. The most useful information isn’t in any single ranked list. It’s in the gaps between lists.
There are many ways to see and identify these gaps.
The “Alternatives” Lens
Cursor sits atop the AI coding IDE leaderboard. GitHub Copilot and Windsurf are virtually tied for second. But ask for alternatives to each of these three, one at a time, and a different name keeps appearing.
Tabnine, ranked #7 in the core category, shows up in 100% of alternative searches across all three tools. Nothing else does. Tabnine isn’t winning the headline race. When it comes to the question developers ask when they’re already unhappy with an incumbent, it’s the most consistent answer in the category. That’s a meaningful position — and it’s invisible from the leaderboard alone.
The “Model Delta” Lens
Comparing results across models reveals editorial fingerprints: When we looked at AWS competitors, Azure was mentioned in every OpenAI response and zero Gemini responses.

“AWS competitors” by LLM provider
Whatever the cause, the implication is real: the model your buyer uses will change which competitors appear beside you.
The “Positioning Gap” Lens
Finally, compare what you say your product does with what LLMs say it does. Your positioning implies a narrative. If you’ve communicated it well, LLMs should agree. You’ll rank well in the use cases you emphasize, weaker showings in the ones you don’t.
Reality is usually messier. When we analyze LLM competitive data with clients, we frequently uncover gaps and opportunities they wouldn’t see otherwise. That popular developer use case that your product supports might not show up in LLMs if you’ve rarely covered it in marketing and documentation sites.
Takeaway: Look at the differences between rankings. Each one points to a different gap, and each is an opportunity to fill.
The right questions naturally expose exactly where to focus.
Know Exactly Where You Rank
That developer with the broken dashboard got a recommendation right where they were building. Whether a product appeared on that list depended on decisions made months earlier. Whether LLMs recommend you depends on what you write, how developers talk about you, and which use cases your content covers.
None of that is visible from a single test query. You can only see it when you look at the pattern across models and with the real developer use cases in their language.
Three ways to put this into practice:
- One query is a data point. Trust the pattern across many.
- The category leaderboard is a starting point. Slice by persona and use case.
- The most useful signal lives in the gaps. Find the ones worth filling.
LLM Rank is the public, free version of this work. Browse it, find your category, and see where you stand. It’s updated monthly, so you’ll always be able to see how LLMs respond to developer questions.
If you want the slices that aren’t on the public site yet — your persona cuts, your alternative queries, the gap between your positioning and what LLMs say — that’s what our LLM Reach Assessment is for. Stop guessing where you stand in AI. Find out exactly.
