Hinglish vs English AI search: same buying question, different brands
Asking the same buying question in Hinglish (Hindi and English typed in English letters) instead of English changes which brands AI names on Gemini and Perplexity, and barely changes ChatGPT. Gemini cited sources in 90.0% of English answers and 41.3% of Hinglish answers. We measured this across 480 answers collected from India, and every response, script and seed is public.

We wrote ten buying questions that name no brand (five skincare, five fashion). Each has an English version and a Hinglish version, and an independent reviewer checked that both ask the same thing. Hinglish here means Hindi and English typed in English letters. Each version ran eight times on ChatGPT, Gemini and Perplexity, geolocated to India, inside a single 57-minute window. That is 480 responses, collected with zero permanent failures for $2.16.
Prompt language changes what AI recommends, and the size of the change depends on the engine far more than on the product category. Measured against each engine's own rerun-to-rerun variation, the language effect on recommended brands is 2.2 points on ChatGPT, 7.8 points on Gemini and 23.4 points on Perplexity. Citations diverge further: Gemini attaches sources to 90.0% of English answers and 41.3% of Hinglish ones. Individual brands swing by up to 17.5 percentage points of visibility between the two languages.
1Why did we run this study?
Hinglish is how many Indians type online: romanized Hindi (Hindi words typed in English letters) mixed with English words, such as "oily skin ke liye konsa face wash kharidna chahiye". Most published research on generative engine optimization (GEO, the work of getting a brand named in AI answers) tests English prompts only.
That gap matters to anyone selling in India. If engines name different brands for a Hinglish version of the same buying question, then English-only tracking describes answers that Hinglish users never see. In the 70 sources we reviewed before collecting data, we found no controlled comparison of English and code-mixed buying prompts (questions that switch between two languages in one sentence), and no GEO measurement run from India.
2How does the study work?
The core problem with measuring an AI engine is that it does not give the same answer twice. Ask the identical question again and the brand list shuffles. English and Hinglish answers will always differ somewhat. The test is whether they differ by more than the engine differs from itself.
Every result below is built on that comparison. We ran each prompt eight times per language, measured how much two same-language reruns overlap (the engine's own noise floor), then measured how much an English answer overlaps a Hinglish one. The difference between those two numbers is the Jaccard gap: how much the brand list changed, compared with how much it changes when you simply ask again. If it is zero, language does not matter.
| Factor | Value |
|---|---|
| Prompts | 10 buying prompts that name no brand, 5 skincare and 5 fashion, frozen before collection |
| Languages | English and romanized Hinglish, pairs independently reviewed for equivalence |
| Engines | ChatGPT with web search and Gemini, both captured from the consumer web apps; Perplexity Sonar (its API answer model) with live web search |
| Market | India-geolocated on every call |
| Repetitions | 8 per prompt per language per engine |
| Responses | 480, zero permanent failures, $2.16 total cost |
| Window | One 57-minute window, randomized interleaved order, fixed seed |
| Statistics | 95% ranges on every rate (Wilson intervals); p-values from checking all 1,024 ways the 10 per-prompt gaps could flip sign by chance (exact permutation test); 95% ranges on each gap from resampling the 10 prompts repeatedly (bootstrap) |
No brand name appears in any prompt, so this measures which brands the engine brings up on its own. The Hinglish versions mix the two languages the way Indian users type them (category nouns in English, function words in romanized Hindi), with no literal translation. Three of the ten pairs failed an independent review and were rewritten before any data was analyzed: one of them lacked an India anchor in its English version, which would have mixed up language with market. That review log is published with the data.
3How much does the brand list change on each engine?
Beyond normal rerun noise, the brand list changed 2.2 points on ChatGPT, 7.8 on Gemini and 23.4 on Perplexity. All three gaps clear the one-sided test we fixed before collection, and the largest is about ten times the smallest. The engine you measure matters more than the category you measure.
| Engine | Rerun baseline | Cross-language | Gap | 95% CI | p (one-sided) |
|---|---|---|---|---|---|
| ChatGPT † | 54.7% | 52.5% | 2.2 pts | 0.3 to 4.6 | .046 |
| Gemini | 51.2% | 43.3% | 7.8 pts | 2.6 to 13.9 | .0098 |
| Perplexity | 68.3% | 44.9% | 23.4 pts | 12.9 to 35.1 | .0010 |
Perplexity is the striking case. It is at once the most self-consistent engine, with the highest rerun overlap of the three at 68.3%, and the most language-sensitive. Its cross-language drop cannot be dismissed as randomness, because its randomness is the lowest in the set.
4How far does Gemini's citation rate drop in Hinglish?
The starkest number in the study has nothing to do with brands. Gemini attached at least one citation to 90.0% of its English answers (72 of 80) and 41.3% of its Hinglish answers (33 of 80). Mean citation count falls from 9.4 to 5.2, and the Hinglish answers are about a third shorter (4,609 down to 3,099 characters on average). The study records the drop. It does not establish the cause, because the design cannot separate how Gemini searches from how it writes.
ChatGPT and Perplexity cite on every answer in both languages, so the drop is specific to Gemini. On the same prompts, in the same window, Gemini gave Hinglish users a shorter answer with fewer sources. It also answered 8 of 80 romanized Hinglish prompts in Devanagari, the script Hindi is usually written in, which the user did not type.
Where engines do cite, the source mix in Indian AI shopping answers looks nothing like a classic search results page. Counting every appearance across the 4,640 citations in the corpus, youtube.com leads with 534 and reddit.com is second with 191, ahead of nykaa.com (176), and a long tail of small Indian review blogs out-cites most mainstream publishers. Counted instead by how many answers cite a domain at all, reddit.com leads with 106 of 480 and youtube.com follows with 100, so the ranking depends on the basis and we state which one we mean. Our guide on how to get cited by AI covers how to show up on the kinds of sources engines cite in India.
5Which brands gain and lose in Hinglish?
Aggregate overlap statistics are the rigorous part, but the practical question for a brand is simpler: does my visibility change? On the language-sensitive engines, by up to 17.5 points.
A pattern worth naming, though we did not test it as a hypothesis and it should be treated as an observation: on Perplexity the brands losing ground in Hinglish skew international or premium, while the gainers skew mass-market Indian direct-to-consumer. The India-origin share of recommended brands rises in the Hinglish arm on Gemini (75.6% to 82%) and Perplexity (63.7% to 66.3%), and stays flat on ChatGPT (74.4% to 74.7%).
Which brand gets named first also becomes less predictable across languages. On Gemini, two same-language reruns agree on the first brand 63.2% of the time, but an English and a Hinglish answer agree only 42.2% of the time. On Perplexity the same figures are 69.5% and 49.8%.
6What should brands selling in India do with this?
Track English and Hinglish separately. If your AI visibility tracking asks English questions only, then on Gemini and Perplexity you are measuring a different brand list from the one Hinglish users see. Our guide on how to measure AI visibility explains how to set up both languages and read them side by side.
Measure each engine on its own. In this study ChatGPT barely moved while Perplexity moved 23.4 points, so an average across engines would hide the change on the engine where it matters most.
Depra tracks Hinglish prompts next to English ones on ChatGPT, Gemini, Perplexity and Google AI Overviews, asked from inside India, and your report shows English and Hinglish visibility side by side. If you are comparing tools, this list of AI visibility tools for D2C brands notes which ones track Hinglish. Plans and limits are on the pricing page.
7How was the study verified?
The full frozen methodology, including the deviations log recording every change made after the design was fixed, is published in the repository. Four verification checks are built into the study, and each of them changed it:
- Prompt review before analysis. An independent reviewer audited all ten pairs for naturalness, equivalence and purchase intent. Three failed and were rewritten, with affected responses discarded and recollected.
- Adversarial recomputation. Six independent automated verification runs (AI agents working from raw data, without our analysis code) recomputed the headline numbers with their own code. All six reproduced them. None refuted them.
- Instrument audit. An audit of 26 brand names that double as ordinary words found three defects in our own tooling, including the brand Simple matching the everyday adjective in 35% of its hits. The verification runs caught two more, both answer-language labels. All five were fixed and the analysis re-run before publication.
- Stability rerun. Two prompts were collected again 30 to 60 minutes later. The direction holds on all three engines, with drift ratios of 0.91, 1.05 and 1.04, where 1.0 means the later answers matched as closely as reruns inside the window.
One detail cuts against our own headline, so it is worth stating plainly. Twenty-six brand names in our lexicon collide with ordinary English or Hindi words, such as the brand Bata against the Hinglish phrase "bata do". We match those case-sensitively. If we relaxed that rule, the ChatGPT gap would inflate from 2.2 to 13.7 points through Hinglish false positives. Any error in our instrument makes the effect look smaller.
8What are the limitations?
Ten prompts in two categories. These results apply to unbranded skincare and fashion buying questions only. One fieldwork window, so this is a snapshot and engines drift over weeks. One Hinglish register, urban and romanized, so Devanagari Hindi and other Indian languages remain untested. Logged-out default surfaces, so signed-in users may see something different. Citations are the sources an engine displays. They do not prove why it named a brand.
The effect sizes differ by engine in a way that fits differences in how each one handles a Hinglish query, but this design does not separate retrieval from generation and cannot attribute the difference to either. That is open work.
This study was designed, funded and run by Depra, which sells AI visibility tracking and therefore has a commercial interest in the conclusion that language matters. We cannot remove that interest, so we tried to make it irrelevant. The prompts were frozen before collection, the statistical tests were fixed before analysis, four verification checks are logged in the repository, every correction they forced is recorded in the deviations log, and the entire corpus is public so a skeptic can recompute every number. The finding we would most like to be true, a large ChatGPT effect, is the one we report as fragile.
9How can you reproduce it?
Everything needed to check our arithmetic is in the public repository: all 480 raw responses with their complete payloads, the frozen prompt set, the brand lexicon, the analysis code and the verification logs. The analysis is deterministic and byte-reproducible, so one command confirms a clean recomputation.
git clone https://github.com/ifham001/IndicGEO
cd IndicGEO
node analyze.mjs
shasum -a 256 data/analysis/results.json | cut -d' ' -f1
# 6f9a67de4c11f11b3e157b1f21799fce66161393d24b5a043eb4a1c395de1776Fresh collection will not reproduce that digest, and it should not: engine answers change over time, so a new run is a new measurement. Replications that disagree with us are welcome as issues or pull requests.
10How to cite this study
Code is MIT licensed, data and text are CC BY 4.0. Machine-readable citation metadata is in the repository's CITATION.cff.
Frequently asked questions
Does ChatGPT recommend different brands in Hinglish?
Barely. ChatGPT's brand list changed 2.2 points more between English and Hinglish than between two reruns of the same question. That result is fragile: the two-sided p is .092, 4 of the 10 prompts moved the other way, and the whole effect sits in the skincare prompts. We describe ChatGPT as approximately language-invariant.
Why does Gemini cite fewer sources in Hinglish?
The study measured the drop but does not establish why. Gemini attached at least one source link to 90.0% of English answers (72 of 80) and 41.3% of Hinglish answers (33 of 80). Its Hinglish answers were also about a third shorter and carried 5.2 citations on average against 9.4. The design cannot separate how Gemini searches from how it writes, so the cause is open work.
Which AI engine changes most with Hinglish prompts?
Perplexity. Its brand list changed 23.4 points more across languages than across reruns (one-sided p = .0010), against 7.8 points on Gemini and 2.2 on ChatGPT. Perplexity was also the most consistent engine on reruns, so the gap stands well clear of its own rerun noise.
How was the study run?
We wrote 10 buying questions (5 skincare, 5 fashion) in English and in Hinglish, none naming a brand. Each version ran 8 times on ChatGPT, Gemini and Perplexity, geolocated to India, inside one 57-minute window on 14 Aug 2026. That gives 480 answers. The full methodology lists every control.
Can I check the data?
Yes. Every raw answer, the prompts, the brand list, the analysis code and the verification logs are public on GitHub. Code is MIT licensed and data is CC BY 4.0. One command, node analyze.mjs, recomputes every number on this page.
A 3-day trial tracks your brand on ChatGPT, Gemini, Perplexity and Google AI Overviews, in English and Hinglish.
Start 3 days free