IndicGEO · Study 01 · Original research

Same question, different brands: what happens when Indian shoppers ask AI in Hinglish instead of English

We asked three AI answer engines the same ten buying questions twice: once in English, once in Hinglish. They did not give the same answers. This is the full study, with every response, script and seed published.

Abstract illustration: two identical speech bubbles, one blue and one violet, each fanning out into a different arrangement of shapes.
Fieldwork 14 Aug 2026480 responsesChatGPT · Gemini · PerplexityIndia-geolocatedReport v1.0
Abstract

Ten brand-unaided purchase-intent prompts (five skincare, five fashion), each written as an independently reviewed semantically equivalent English/Hinglish pair, were run eight times per language against ChatGPT, Gemini and Perplexity, geolocated to India, inside a single 57-minute window. That is 480 responses, collected with zero permanent failures for $2.16.

Prompt language changes what AI recommends, and the size of the change depends on the engine far more than on the product category. Measured against each engine's own rerun-to-rerun variability, the language effect on recommended brand sets is 2.2 points on ChatGPT, 7.8 points on Gemini and 23.4 points on Perplexity. Citations diverge harder: Gemini attaches sources to 90% of English answers but only 41.3% of Hinglish ones. Individual brands swing by up to 17.5 percentage points of visibility between the two languages.

90% → 41.3%
Gemini answers carrying a citation
English 72/80, Hinglish 33/80
23.4 pts
Perplexity brand-set language gap
vs its own rerun noise, p = .001
17.5 pts
Largest single brand swing (The Derma Co)
2/80 English → 16/80 Hinglish

1Why we ran this

A large share of Indian AI-assistant usage is code-mixed: people type romanized Hinglish, not English and not Devanagari. Almost every AI-visibility tool on the market, and almost all published research on generative engine optimization, measures English prompts only.

That gap matters commercially and scientifically. If engines name different brands for a Hinglish version of the same buying question, then English-only tracking describes a market that a large share of Indian users never see. Before this study, no controlled English-versus-code-mixed comparison of buying prompts existed, no GEO measurement had been geolocated to India, and nothing had tested romanized code-mix at all.

2How the study works

The core problem with measuring anything about an AI engine is that it does not give the same answer twice. Ask the identical question again and the brand list shuffles. So the interesting question is not "do English and Hinglish answers differ", because they always will. It is whether they differ by more than the engine differs from itself.

Every result below is built on that comparison. We ran each prompt eight times per language, measured how much two same-language reruns overlap (the engine's own noise floor), then measured how much an English answer overlaps a Hinglish one. The gap between those two numbers is the language effect. If it is zero, language does not matter.

FactorValue
Prompts10 brand-unaided buying prompts, 5 skincare and 5 fashion, frozen before collection
LanguagesEnglish and romanized Hinglish, pairs independently reviewed for equivalence
EnginesChatGPT and Gemini as consumer-UI captures, Perplexity Sonar with live web search
MarketIndia-geolocated on every call
Repetitions8 per prompt per language per engine
Responses480, zero permanent failures, $2.16 total cost
WindowOne 57-minute window, randomized interleaved order, fixed seed
StatisticsWilson intervals, exact sign-flip permutation tests, clustered bootstrap

No brand name appears in any prompt, so this measures unaided recall. The Hinglish versions are authentic code-mix as Indian users actually type (category nouns in English, function words in romanized Hindi), not machine translation. Three of the ten pairs failed an independent pre-analysis review and were rewritten before any data was analyzed: one of them lacked an India anchor in its English version, which would have confounded language with market. That review log is published with the data.

3What we found

The language effect is real on all three engines, but it spans an order of magnitude between them. The engine you measure matters more than the category you measure.

Recommended-brand overlap: identical reruns vs across languages
Same language, rerun to rerunEnglish vs Hinglish
ChatGPT
54.7%
52.5%gap 2.2 pts, p = .046
Gemini
51.2%
43.3%gap 7.8 pts, p = .010
Perplexity
68.3%
44.9%gap 23.4 pts, p = .001
Mean pairwise Jaccard similarity of the set of brands named in each answer. Baseline pairs every same-language repetition of a prompt (56 pairs per prompt per engine); cross-language pairs every English answer with every Hinglish answer of the same prompt (64 pairs). A gap means language changes the recommendation set by more than the engine changes on its own.
EngineRerun baselineCross-languageGap95% CIp
ChatGPT †54.7%52.5%2.2 pts0.3 to 4.6.0459
Gemini51.2%43.3%7.8 pts2.6 to 13.9.0098
Perplexity68.3%44.9%23.4 pts12.9 to 35.1.001
† ChatGPT is the fragile result and we flag it rather than lean on it: the two-sided p is .092, four of its ten per-prompt deltas are negative, and the effect is carried entirely by the skincare half (skincare 4.45 points, p = .063; fashion 0.01 points, p = .469). We describe ChatGPT as approximately language-invariant.

Perplexity is the striking case. It is simultaneously the most self-consistent engine, with the highest rerun overlap of the three at 68.3%, and the most language-sensitive. Its cross-language drop cannot be dismissed as randomness, because its randomness is the lowest in the set.

4The Gemini citation collapse

The single starkest number in the study has nothing to do with brands. Gemini attached at least one citation to 90% of its English answers (72 of 80) and only 41.3% of its Hinglish answers (33 of 80). Mean citation count falls from 9.4 to 5.2, and the Hinglish answers are about a third shorter (4,609 down to 3,099 characters on average).

Share of answers carrying at least one citation
English promptHinglish prompt
ChatGPT
100% (80/80)
100% (80/80)
Gemini
90% (72/80)
41.3% (33/80)
Perplexity
100% (80/80)
100% (80/80)
Counts are out of 80 responses per engine per language. ChatGPT and Perplexity cite on every answer in both languages. Gemini's Hinglish answers also carry fewer citations on average (9.4 falling to 5.2) and are about a third shorter.

ChatGPT and Perplexity cite on every answer in both languages, so this is specific to Gemini rather than a property of code-mixed queries in general. On the same prompts, in the same window, Gemini gave code-mixed users a shorter, less-sourced answer. It also answered 8 of 80 romanized Hinglish prompts in Devanagari script, a register the user did not type.

Where engines do cite, the source mix in Indian AI shopping answers looks nothing like a classic search results page. YouTube is the most-cited domain in the whole corpus, Reddit is fourth, and a long tail of small Indian review blogs out-cites most mainstream publishers.

5Which brands win and lose

Aggregate overlap statistics are the rigorous part, but the practical question for a brand is simpler: does my visibility change? On the language-sensitive engines, substantially.

Brands that gain or lose visibility in Hinglish: Perplexity
Brand
Change
The Derma Co
+17.5
Taneira
15
Plum
13.7
La Roche-Posay
12.5
Deconstruct
+12.5
Minimalist
+11.3
Vastranand
10
Garnier
+10
Change in the share of 80 responses per language that mention each brand, for the ten brands that move most. Bars right of the centre line mean the brand is named more often when the question is asked in Hinglish. Percentages are shown with the underlying counts because a single response is 1.25 percentage points.

A pattern worth naming, though we did not test it as a hypothesis and it should be treated as an observation: on Perplexity the brands losing ground in Hinglish skew international or premium, while the gainers skew mass-market Indian direct-to-consumer. The India-origin share of recommended brands rises in the Hinglish arm on Gemini (75.6% to 82%) and Perplexity (63.7% to 66.3%), and stays flat on ChatGPT (74.4% to 74.7%).

Brands that gain or lose visibility in Hinglish: Gemini
Brand
Change
Dot & Key
+11.3
CeraVe
10
Allen Solly
+10
The Derma Co
8.8
Zodiac
8.8
Sparx
+8.8
Change in the share of 80 responses per language that mention each brand, for the ten brands that move most. Bars right of the centre line mean the brand is named more often when the question is asked in Hinglish. Percentages are shown with the underlying counts because a single response is 1.25 percentage points.

Which brand gets named first also becomes less predictable across languages. On Gemini, two same-language reruns agree on the first brand 63.2% of the time, but an English and a Hinglish answer agree only 42.2% of the time. On Perplexity the same figures are 69.5% and 49.8%.

6Methodology and verification

The full frozen methodology, including the deviations log recording every change made after the design was fixed, is published in the repository. Four verification steps are built into the study, and each of them changed it:

  • Prompt review before analysis. An independent reviewer audited all ten pairs for naturalness, equivalence and purchase intent. Three failed and were rewritten, with affected responses discarded and recollected.
  • Adversarial recomputation. Six independent agents recomputed the headline numbers straight from the raw data using their own code, with no access to our analysis pipeline. All six reproduced; none refuted.
  • Instrument audit. That same pass found five measurement defects in our own tooling, including a brand name that was matching as a generic English adjective in 35% of its hits. All were fixed and the analysis re-run before publication.
  • Stability rerun. Two prompts were collected again 30 to 60 minutes later. The direction holds on all three engines, with drift ratios of 0.91, 1.05 and 1.04.

One detail cuts against our own headline, so it is worth stating plainly. Twenty-six brand names in our lexicon collide with ordinary English or Hindi words, such as the brand Bata against the Hinglish phrase "bata do". We match those case-sensitively. If we relaxed that rule, the ChatGPT gap would inflate from 2.2 to 13.7 points through Hinglish false positives. Our instrument understates the effect rather than manufacturing it.

7Limitations

Ten prompts in two categories. These results describe this class of prompt, not Indian buying queries at large. One fieldwork window, so this is a snapshot and engines drift over weeks. One Hinglish register, urban and romanized, so Devanagari Hindi and other Indian languages remain untested. Logged-out default surfaces, so signed-in users may see something different. Citations are the sources an engine displays, not verified evidence for why it named a brand.

The effect-size hierarchy across engines is consistent with differences in how each one handles a code-mixed query, but this design does not separate retrieval from generation and cannot attribute the difference to either. That is open work.

Funding and interest

This study was designed, funded and run by Depra, which sells AI-visibility tracking and therefore has a commercial interest in the conclusion that language matters. We cannot remove that interest, so we tried to make it irrelevant. The prompts were frozen before collection, the statistical tests were fixed before analysis, three rounds of independent verification are logged in the repository, every correction those rounds forced is recorded in the deviations log, and the entire corpus is public so a skeptic can recompute every number. The finding we would most like to be true, a large ChatGPT effect, is the one we report as fragile.

8Reproduce it yourself

Everything needed to check our arithmetic is in the public repository: all 480 raw responses with their complete API payloads, the frozen prompt set, the brand lexicon, the analysis code and the verification logs. The analysis is deterministic and byte-reproducible, so one command confirms a clean recomputation.

git clone https://github.com/ifham001/IndicGEO
cd IndicGEO
node analyze.mjs
shasum -a 256 data/analysis/results.json | cut -d' ' -f1
# a1afac1d716bf845fdee2d31941b672c1e728d44e459eac13c854740a12203e2

Fresh collection will not reproduce that digest, and it should not: engine answers change over time, so a new run is a new measurement rather than a replication. Replications that disagree with us are welcome as issues or pull requests.

9How to cite

Depra Research (2026). IndicGEO Study 01: Same question, different brands. English versus Hinglish buying prompts across ChatGPT, Gemini, and Perplexity in India (Version 1.0.0) [Data set]. https://github.com/ifham001/IndicGEO

Code is MIT licensed, data and text are CC BY 4.0. Machine-readable citation metadata is in the repository's CITATION.cff.

Is your brand visible in both languages?

Depra tracks how AI answer engines name brands in Indian markets, in English and in Hinglish. The free audit runs your brand through the same engines used in this study.

Run my free audit