Yes, AI engines recommend different brands when the same buying question is asked in Hinglish, and how much depends on the engine. In our study of 480 AI answers collected from India, Perplexity's brand list changed 23.4 points beyond normal rerun noise, Gemini's changed 7.8 points and ChatGPT's changed only 2.2 points.
Hinglish here means Hindi and English mixed and typed in English letters, like "konsa face wash kharidna chahiye". Rerun noise is how much an engine's brand list changes when you simply ask the identical question again. The second headline is about sources: Gemini showed at least one source link under 90.0% of English answers and under only 41.3% of Hinglish answers.
This post is the plain readout for brand and marketing teams. The Hinglish vs English AI shopping study page is the record of every table, interval and test, and the raw data is public in the IndicGEO repository on GitHub.
How we tested the same question in two languages
We wrote 10 buying questions: 5 about skincare and 5 about fashion. These are the prompts, meaning the exact text typed into each AI engine, written the way a buyer would type it. None of them named a brand. Each question had an English version and a Hinglish version with the same meaning, the same budget and the same India anchor. An independent reviewer checked every pair before analysis, and 3 fashion pairs were rewritten after failing that review.
One pair, word for word:
- English: "Which face wash should I buy for oily skin in India under ₹500?"
- Hinglish: "India mein oily skin ke liye 500 rupees ke andar konsa face wash kharidna chahiye?"
We asked each question 8 times in each language on three engines:
- ChatGPT, captured from the chatgpt.com web app with web search switched on.
- Gemini, captured from the gemini.google.com web app.
- Perplexity, through Sonar with live web search switched on. Sonar is the answer model Perplexity offers to developers, reached through its API instead of the perplexity.ai app.
Every request was located in India and started fresh, with no chat history and no logged-in account. All 480 answers were collected on 14 August 2026 inside one 57-minute window, between 12:25 and 13:22 UTC. The order was shuffled, so neither language got a different time of day.
Brands were counted by matching each answer against a brand list built from all 480 answers and checked by hand. Some brand names are also everyday Hindi words. "Bata" is a shoe brand and also the Hinglish verb for "tell". Those names counted only when capitalised as a brand, so Hinglish answers were not over-counted.
Why rerun noise is the baseline
An AI engine rarely gives the same answer twice. English and Hinglish answers will always differ a little, even if language has no effect. So we measured two things per engine:
- How much the brand list overlaps between reruns of the same question in the same language.
- How much it overlaps between the English and the Hinglish version of that question.
The overlap score is the number of brands two answers share, divided by all brands either answer named. The language effect is the gap between the two overlaps. A gap of zero means language changes nothing beyond chance.
This follows a point made in Don't Measure Once, a 2026 paper by Schulte, Bleeker and Kaufmann: AI answers vary across runs, prompts and time, so one answer is an unreliable measure of a brand's visibility.
Which AI engine changes its brand recommendations most?
Perplexity, by a wide margin.
| Engine | Brand overlap between reruns | Brand overlap, English vs Hinglish | Language effect | 95% range |
|---|---|---|---|---|
| ChatGPT | 54.7% | 52.5% | 2.2 points | 0.3 to 4.6 |
| Gemini | 51.2% | 43.3% | 7.8 points | 2.6 to 13.9 |
| Perplexity | 68.3% | 44.9% | 23.4 points | 12.9 to 35.1 |
Basis: 480 answers, 80 per engine and language (10 prompts x 8 runs), India-located, 14 August 2026. The 95% range is the span the true language effect most likely falls in, given that the study used only 10 prompts. It comes from resampling the 10 prompts 10,000 times.
Perplexity is the most consistent engine on reruns, at 68.3% brand overlap, and also the most sensitive to language. When the most self-consistent engine drops to 44.9% overlap across languages, chance is a poor explanation. Its choice of which brand to name first also shifted: the first brand matched in 69.5% of rerun pairs and 49.8% of English and Hinglish pairs.
Gemini sits in the middle at 7.8 points. Its first-named brand matched in 63.2% of rerun pairs and 42.2% of cross-language pairs.
ChatGPT barely moves. The study fixed its test before looking at the data: a one-sided test, which only asks whether English and Hinglish answers overlap less than reruns do. The 2.2-point gap passed it (p = .046, meaning a gap this size would show up by chance less than 5 times in 100 if language made no difference). It is still the weakest result in the study. The whole gap came from the skincare prompts and 4 of the 10 prompts went the other way. A two-sided test, which also allows for Hinglish overlapping more, would not pass (p = .092). The fair reading is that Hinglish ChatGPT answers named close to the same brands as English ones for these questions.
The sources behind the answers follow the same order. The overlap in cited websites between languages fell 3.1 points on ChatGPT, 22.3 points on Gemini and 32.4 points on Perplexity, each measured against the same rerun baseline. If you track several engines, why AI engines disagree about your brand covers what drives those differences.
Which brands appeared or vanished in Hinglish answers?
The engine-level gaps are averages. For single brands, the swings on Perplexity were large. Each figure below is the share of one engine's 80 answers in one language that named the brand.
Perplexity
| Brand | English answers | Hinglish answers | Change |
|---|---|---|---|
| The Derma Co | 2.5% (2 of 80) | 20.0% (16 of 80) | +17.5 points |
| Deconstruct | 3.8% (3 of 80) | 16.3% (13 of 80) | +12.5 points |
| Minimalist | 22.5% (18 of 80) | 33.8% (27 of 80) | +11.3 points |
| Plum | 28.7% (23 of 80) | 15.0% (12 of 80) | -13.7 points |
| La Roche-Posay | 13.8% (11 of 80) | 1.3% (1 of 80) | -12.5 points |
| Taneira | 15.0% (12 of 80) | 0.0% (0 of 80) | -15.0 points |
Gemini
| Brand | English answers | Hinglish answers | Change |
|---|---|---|---|
| Dot & Key | 10.0% (8 of 80) | 21.3% (17 of 80) | +11.3 points |
| Allen Solly | 5.0% (4 of 80) | 15.0% (12 of 80) | +10.0 points |
| Minimalist | 45.0% (36 of 80) | 36.3% (29 of 80) | -8.7 points |
| CeraVe | 10.0% (8 of 80) | 0.0% (0 of 80) | -10.0 points |
On ChatGPT, no brand moved by more than 10 points. The largest change was Deconstruct, from 10.0% to 18.8% of answers.
One pattern stood out on Perplexity. Brands that lost ground in Hinglish leaned international or premium, such as La Roche-Posay and Taneira, a premium Tata brand. Brands that gained leaned toward mass-market Indian D2C brands, such as The Derma Co, Deconstruct and Minimalist. D2C, short for direct-to-consumer, describes brands built to sell online straight to shoppers. The share of named brands that are Indian rose in Hinglish on Gemini, from 75.6% to 82.0%, and on Perplexity, from 63.7% to 66.3%. On ChatGPT it stayed flat, 74.4% to 74.7%. We observed this tendency and did not test it as a hypothesis.
Treat these as observations about AI shopping recommendations in India, from 10 questions on one day. They are not endorsements and they say nothing about which product is better. The same brands could score very differently on another set of questions.
Gemini cites far fewer sources in Hinglish
A citation is a source link an engine shows with its answer. Gemini attached at least one to 72 of 80 English answers (90.0%) and to only 33 of 80 Hinglish answers (41.3%). Its average number of links per answer fell from 9.4 to 5.2. Its Hinglish answers were also 33% shorter, averaging 3,099 characters against 4,609 in English.
ChatGPT and Perplexity attached sources to every answer in both languages, 80 of 80 each. What changed on those two engines was which websites they cited.
The study did not test why Gemini behaves this way. What it recorded: 47 of 80 Gemini Hinglish answers showed no source links at all, against 8 of 80 English answers.
Do AI engines answer Hinglish questions in Hinglish?
Almost always.
| Engine | Asked in English | Asked in Hinglish |
|---|---|---|
| ChatGPT | 80 of 80 in English | 80 of 80 in Hinglish |
| Gemini | 80 of 80 in English | 71 in Hinglish, 8 in Devanagari, 1 in English |
| Perplexity | 80 of 80 in English | 80 of 80 in Hinglish |
Devanagari is the script Hindi is usually printed in. Gemini used it for 8 of its 80 Hinglish answers, a script the person asking did not type. The language labels were checked against human judgement on a sample of 30 answers and agreed on all 30.
So a Hinglish buyer reads a Hinglish answer. On Perplexity and Gemini that answer also links to a different set of websites and names a different set of brands than the English one.
What does this mean for a brand selling in India?
If you track English prompts only, you are measuring what English buyers see and leaving out Hinglish AI search. On Perplexity and Gemini, Hinglish buyers see a different brand list. Five steps follow from the data:
- Write each buying question twice. Keep the budget, category and intent identical in English and Hinglish. Our guide to writing Hinglish tracking prompts covers the wording.
- Read each engine on its own. The language effect ran from 2.2 points on ChatGPT to 23.4 points on Perplexity. An average across engines hides the one where you lose most.
- Ask more than once before trusting a number. Rerun overlap was 51% to 68% for brands, so a single answer tells you little.
- Compare your English and Hinglish share with your top competitors. A split like The Derma Co's 2.5% against 20.0% on Perplexity is the kind of gap to look for.
- Check which websites get cited in each language. Cited-site overlap between languages fell 32.4 points on Perplexity and 22.3 points on Gemini, so the pages that earn you mentions in English may not be the ones Hinglish answers link to.
Depra runs this as a product. Hinglish AI visibility tracking is on every plan: Depra suggests Hinglish versions of your buying questions when India is one of your markets, detects the language of every answer, and your report shows an English vs Hinglish table. ChatGPT, Gemini and Google AI Overviews, the AI-written summary Google shows above some search results, are checked daily. The Perplexity visibility tracker runs weekly. Plans start at ₹1,999 a month plus 18% GST, listed on the pricing page. For the basics of reading these numbers, see how to track brand mentions in AI answers.
Limits of the study
- 10 prompts in 2 categories. The results describe brand-free buying questions about skincare and fashion in India. Other categories may behave differently.
- One fieldwork window. Engines change week to week, so this is a snapshot from 14 August 2026. A rerun of 2 prompts 30 to 60 minutes later pointed the same way on all three engines, which covers drift within the hour only.
- One style of Hinglish. Urban, typed in English letters, with 43% to 69% Hindi words per prompt. Devanagari Hindi and other Indian languages were out of scope.
- Logged-out sessions. Signed-in users with a chat history may see different answers.
- Displayed sources only. A citation records what the engine linked, and does not prove why a brand was named.
- ChatGPT is the weakest finding. Its 2.2-point gap rests on the skincare prompts alone.
- Funding interest. Depra ran and paid for the study and sells AI visibility tracking. To offset that, every raw answer, the analysis code, the brand list and the method, fixed before any analysis, are public in the IndicGEO repository, and one command recomputes every table in this post.
To see your own brand's English and Hinglish numbers across ChatGPT, Gemini, Perplexity and AI Overviews, start 7 days free on any plan, no card. Your first scan starts at signup.
