Same question, different brands: what happens when Indian shoppers ask AI in Hinglish instead of English
We asked three AI answer engines the same ten buying questions twice: once in English, once in Hinglish. They did not give the same answers. This is the full study, with every response, script and seed published.

Ten brand-unaided purchase-intent prompts (five skincare, five fashion), each written as an independently reviewed semantically equivalent English/Hinglish pair, were run eight times per language against ChatGPT, Gemini and Perplexity, geolocated to India, inside a single 57-minute window. That is 480 responses, collected with zero permanent failures for $2.16.
Prompt language changes what AI recommends, and the size of the change depends on the engine far more than on the product category. Measured against each engine's own rerun-to-rerun variability, the language effect on recommended brand sets is 2.2 points on ChatGPT, 7.8 points on Gemini and 23.4 points on Perplexity. Citations diverge harder: Gemini attaches sources to 90% of English answers but only 41.3% of Hinglish ones. Individual brands swing by up to 17.5 percentage points of visibility between the two languages.
1Why we ran this
A large share of Indian AI-assistant usage is code-mixed: people type romanized Hinglish, not English and not Devanagari. Almost every AI-visibility tool on the market, and almost all published research on generative engine optimization, measures English prompts only.
That gap matters commercially and scientifically. If engines name different brands for a Hinglish version of the same buying question, then English-only tracking describes a market that a large share of Indian users never see. Before this study, no controlled English-versus-code-mixed comparison of buying prompts existed, no GEO measurement had been geolocated to India, and nothing had tested romanized code-mix at all.
2How the study works
The core problem with measuring anything about an AI engine is that it does not give the same answer twice. Ask the identical question again and the brand list shuffles. So the interesting question is not "do English and Hinglish answers differ", because they always will. It is whether they differ by more than the engine differs from itself.
Every result below is built on that comparison. We ran each prompt eight times per language, measured how much two same-language reruns overlap (the engine's own noise floor), then measured how much an English answer overlaps a Hinglish one. The gap between those two numbers is the language effect. If it is zero, language does not matter.
| Factor | Value |
|---|---|
| Prompts | 10 brand-unaided buying prompts, 5 skincare and 5 fashion, frozen before collection |
| Languages | English and romanized Hinglish, pairs independently reviewed for equivalence |
| Engines | ChatGPT and Gemini as consumer-UI captures, Perplexity Sonar with live web search |
| Market | India-geolocated on every call |
| Repetitions | 8 per prompt per language per engine |
| Responses | 480, zero permanent failures, $2.16 total cost |
| Window | One 57-minute window, randomized interleaved order, fixed seed |
| Statistics | Wilson intervals, exact sign-flip permutation tests, clustered bootstrap |
No brand name appears in any prompt, so this measures unaided recall. The Hinglish versions are authentic code-mix as Indian users actually type (category nouns in English, function words in romanized Hindi), not machine translation. Three of the ten pairs failed an independent pre-analysis review and were rewritten before any data was analyzed: one of them lacked an India anchor in its English version, which would have confounded language with market. That review log is published with the data.
3What we found
The language effect is real on all three engines, but it spans an order of magnitude between them. The engine you measure matters more than the category you measure.
| Engine | Rerun baseline | Cross-language | Gap | 95% CI | p |
|---|---|---|---|---|---|
| ChatGPT † | 54.7% | 52.5% | 2.2 pts | 0.3 to 4.6 | .0459 |
| Gemini | 51.2% | 43.3% | 7.8 pts | 2.6 to 13.9 | .0098 |
| Perplexity | 68.3% | 44.9% | 23.4 pts | 12.9 to 35.1 | .001 |
Perplexity is the striking case. It is simultaneously the most self-consistent engine, with the highest rerun overlap of the three at 68.3%, and the most language-sensitive. Its cross-language drop cannot be dismissed as randomness, because its randomness is the lowest in the set.
4The Gemini citation collapse
The single starkest number in the study has nothing to do with brands. Gemini attached at least one citation to 90% of its English answers (72 of 80) and only 41.3% of its Hinglish answers (33 of 80). Mean citation count falls from 9.4 to 5.2, and the Hinglish answers are about a third shorter (4,609 down to 3,099 characters on average).
ChatGPT and Perplexity cite on every answer in both languages, so this is specific to Gemini rather than a property of code-mixed queries in general. On the same prompts, in the same window, Gemini gave code-mixed users a shorter, less-sourced answer. It also answered 8 of 80 romanized Hinglish prompts in Devanagari script, a register the user did not type.
Where engines do cite, the source mix in Indian AI shopping answers looks nothing like a classic search results page. YouTube is the most-cited domain in the whole corpus, Reddit is fourth, and a long tail of small Indian review blogs out-cites most mainstream publishers.
5Which brands win and lose
Aggregate overlap statistics are the rigorous part, but the practical question for a brand is simpler: does my visibility change? On the language-sensitive engines, substantially.
A pattern worth naming, though we did not test it as a hypothesis and it should be treated as an observation: on Perplexity the brands losing ground in Hinglish skew international or premium, while the gainers skew mass-market Indian direct-to-consumer. The India-origin share of recommended brands rises in the Hinglish arm on Gemini (75.6% to 82%) and Perplexity (63.7% to 66.3%), and stays flat on ChatGPT (74.4% to 74.7%).
Which brand gets named first also becomes less predictable across languages. On Gemini, two same-language reruns agree on the first brand 63.2% of the time, but an English and a Hinglish answer agree only 42.2% of the time. On Perplexity the same figures are 69.5% and 49.8%.
6Methodology and verification
The full frozen methodology, including the deviations log recording every change made after the design was fixed, is published in the repository. Four verification steps are built into the study, and each of them changed it:
- Prompt review before analysis. An independent reviewer audited all ten pairs for naturalness, equivalence and purchase intent. Three failed and were rewritten, with affected responses discarded and recollected.
- Adversarial recomputation. Six independent agents recomputed the headline numbers straight from the raw data using their own code, with no access to our analysis pipeline. All six reproduced; none refuted.
- Instrument audit. That same pass found five measurement defects in our own tooling, including a brand name that was matching as a generic English adjective in 35% of its hits. All were fixed and the analysis re-run before publication.
- Stability rerun. Two prompts were collected again 30 to 60 minutes later. The direction holds on all three engines, with drift ratios of 0.91, 1.05 and 1.04.
One detail cuts against our own headline, so it is worth stating plainly. Twenty-six brand names in our lexicon collide with ordinary English or Hindi words, such as the brand Bata against the Hinglish phrase "bata do". We match those case-sensitively. If we relaxed that rule, the ChatGPT gap would inflate from 2.2 to 13.7 points through Hinglish false positives. Our instrument understates the effect rather than manufacturing it.
7Limitations
Ten prompts in two categories. These results describe this class of prompt, not Indian buying queries at large. One fieldwork window, so this is a snapshot and engines drift over weeks. One Hinglish register, urban and romanized, so Devanagari Hindi and other Indian languages remain untested. Logged-out default surfaces, so signed-in users may see something different. Citations are the sources an engine displays, not verified evidence for why it named a brand.
The effect-size hierarchy across engines is consistent with differences in how each one handles a code-mixed query, but this design does not separate retrieval from generation and cannot attribute the difference to either. That is open work.
This study was designed, funded and run by Depra, which sells AI-visibility tracking and therefore has a commercial interest in the conclusion that language matters. We cannot remove that interest, so we tried to make it irrelevant. The prompts were frozen before collection, the statistical tests were fixed before analysis, three rounds of independent verification are logged in the repository, every correction those rounds forced is recorded in the deviations log, and the entire corpus is public so a skeptic can recompute every number. The finding we would most like to be true, a large ChatGPT effect, is the one we report as fragile.
8Reproduce it yourself
Everything needed to check our arithmetic is in the public repository: all 480 raw responses with their complete API payloads, the frozen prompt set, the brand lexicon, the analysis code and the verification logs. The analysis is deterministic and byte-reproducible, so one command confirms a clean recomputation.
git clone https://github.com/ifham001/IndicGEO
cd IndicGEO
node analyze.mjs
shasum -a 256 data/analysis/results.json | cut -d' ' -f1
# a1afac1d716bf845fdee2d31941b672c1e728d44e459eac13c854740a12203e2Fresh collection will not reproduce that digest, and it should not: engine answers change over time, so a new run is a new measurement rather than a replication. Replications that disagree with us are welcome as issues or pull requests.
9How to cite
Code is MIT licensed, data and text are CC BY 4.0. Machine-readable citation metadata is in the repository's CITATION.cff.
Depra tracks how AI answer engines name brands in Indian markets, in English and in Hinglish. The free audit runs your brand through the same engines used in this study.
Run my free audit