IndicGEO · Study 01 · Original research

Hinglish vs English AI search: same buying question, different brands

Asking the same buying question in Hinglish (Hindi and English typed in English letters) instead of English changes which brands AI names on Gemini and Perplexity, and barely changes ChatGPT. Gemini cited sources in 90.0% of English answers and 41.3% of Hinglish answers. We measured this across 480 answers collected from India, and every response, script and seed is public.

Abstract illustration: two identical speech bubbles, one blue and one violet, each fanning out into a different arrangement of shapes.
By Depra ResearchFieldwork 14 Aug 2026Updated 15 Sep 2026480 responsesChatGPT · Gemini · PerplexityIndia-geolocatedReport v1.0
Abstract

We wrote ten buying questions that name no brand (five skincare, five fashion). Each has an English version and a Hinglish version, and an independent reviewer checked that both ask the same thing. Hinglish here means Hindi and English typed in English letters. Each version ran eight times on ChatGPT, Gemini and Perplexity, geolocated to India, inside a single 57-minute window. That is 480 responses, collected with zero permanent failures for $2.16.

Prompt language changes what AI recommends, and the size of the change depends on the engine far more than on the product category. Measured against each engine's own rerun-to-rerun variation, the language effect on recommended brands is 2.2 points on ChatGPT, 7.8 points on Gemini and 23.4 points on Perplexity. Citations diverge further: Gemini attaches sources to 90.0% of English answers and 41.3% of Hinglish ones. Individual brands swing by up to 17.5 percentage points of visibility between the two languages.

41.3%
Gemini Hinglish citation rate
Down from 90.0% in English (72/80 to 33/80)
23.4 pts
Perplexity brand-list gap
vs its own rerun noise, p = .0010
17.5 pts
Biggest swing: The Derma Co
2/80 English → 16/80 Hinglish

1Why did we run this study?

Hinglish is how many Indians type online: romanized Hindi (Hindi words typed in English letters) mixed with English words, such as "oily skin ke liye konsa face wash kharidna chahiye". Most published research on generative engine optimization (GEO, the work of getting a brand named in AI answers) tests English prompts only.

That gap matters to anyone selling in India. If engines name different brands for a Hinglish version of the same buying question, then English-only tracking describes answers that Hinglish users never see. In the 70 sources we reviewed before collecting data, we found no controlled comparison of English and code-mixed buying prompts (questions that switch between two languages in one sentence), and no GEO measurement run from India.

2How does the study work?

The core problem with measuring an AI engine is that it does not give the same answer twice. Ask the identical question again and the brand list shuffles. English and Hinglish answers will always differ somewhat. The test is whether they differ by more than the engine differs from itself.

Every result below is built on that comparison. We ran each prompt eight times per language, measured how much two same-language reruns overlap (the engine's own noise floor), then measured how much an English answer overlaps a Hinglish one. The difference between those two numbers is the Jaccard gap: how much the brand list changed, compared with how much it changes when you simply ask again. If it is zero, language does not matter.

FactorValue
Prompts10 buying prompts that name no brand, 5 skincare and 5 fashion, frozen before collection
LanguagesEnglish and romanized Hinglish, pairs independently reviewed for equivalence
EnginesChatGPT with web search and Gemini, both captured from the consumer web apps; Perplexity Sonar (its API answer model) with live web search
MarketIndia-geolocated on every call
Repetitions8 per prompt per language per engine
Responses480, zero permanent failures, $2.16 total cost
WindowOne 57-minute window, randomized interleaved order, fixed seed
Statistics95% ranges on every rate (Wilson intervals); p-values from checking all 1,024 ways the 10 per-prompt gaps could flip sign by chance (exact permutation test); 95% ranges on each gap from resampling the 10 prompts repeatedly (bootstrap)

No brand name appears in any prompt, so this measures which brands the engine brings up on its own. The Hinglish versions mix the two languages the way Indian users type them (category nouns in English, function words in romanized Hindi), with no literal translation. Three of the ten pairs failed an independent review and were rewritten before any data was analyzed: one of them lacked an India anchor in its English version, which would have mixed up language with market. That review log is published with the data.

3How much does the brand list change on each engine?

Beyond normal rerun noise, the brand list changed 2.2 points on ChatGPT, 7.8 on Gemini and 23.4 on Perplexity. All three gaps clear the one-sided test we fixed before collection, and the largest is about ten times the smallest. The engine you measure matters more than the category you measure.

Recommended-brand overlap: identical reruns vs across languages
Same language, rerun to rerunEnglish vs Hinglish
ChatGPT
54.7%
52.5%
gap 2.2 pts, p = .046
Gemini
51.2%
43.3%
gap 7.8 pts, p = .0098
Perplexity
68.3%
44.9%
gap 23.4 pts, p = .0010
Mean pairwise Jaccard similarity of the brands named in each answer: brands named in both answers, divided by brands named in either. Baseline pairs every same-language repetition of a prompt (56 pairs per prompt per engine); cross-language pairs every English answer with every Hinglish answer of the same prompt (64 pairs). A gap means language changes the recommendation set by more than the engine changes on its own. p is an exact one-sided test: the chance of a gap this large if language made no difference.
EngineRerun baselineCross-languageGap95% CIp (one-sided)
ChatGPT †54.7%52.5%2.2 pts0.3 to 4.6.046
Gemini51.2%43.3%7.8 pts2.6 to 13.9.0098
Perplexity68.3%44.9%23.4 pts12.9 to 35.1.0010
p is the chance of seeing a gap this large if language made no difference. 95% CI is the range the gap most likely sits in. † ChatGPT is the fragile result, so we flag it: the two-sided p is .092, four of its ten per-prompt deltas are negative, and the effect is carried entirely by the skincare half (skincare 4.45 points, p = .063; fashion 0.01 points, p = .469). We describe ChatGPT as approximately language-invariant.

Perplexity is the striking case. It is at once the most self-consistent engine, with the highest rerun overlap of the three at 68.3%, and the most language-sensitive. Its cross-language drop cannot be dismissed as randomness, because its randomness is the lowest in the set.

4How far does Gemini's citation rate drop in Hinglish?

The starkest number in the study has nothing to do with brands. Gemini attached at least one citation to 90.0% of its English answers (72 of 80) and 41.3% of its Hinglish answers (33 of 80). Mean citation count falls from 9.4 to 5.2, and the Hinglish answers are about a third shorter (4,609 down to 3,099 characters on average). The study records the drop. It does not establish the cause, because the design cannot separate how Gemini searches from how it writes.

Share of answers carrying at least one citation
English promptHinglish prompt
ChatGPT
100% (80/80)
100% (80/80)
Gemini
90% (72/80)
41.3% (33/80)
Perplexity
100% (80/80)
100% (80/80)
Counts are out of 80 responses per engine per language. ChatGPT and Perplexity cite on every answer in both languages. Gemini's Hinglish answers also carry fewer citations on average (9.4 falling to 5.2) and are about a third shorter.

ChatGPT and Perplexity cite on every answer in both languages, so the drop is specific to Gemini. On the same prompts, in the same window, Gemini gave Hinglish users a shorter answer with fewer sources. It also answered 8 of 80 romanized Hinglish prompts in Devanagari, the script Hindi is usually written in, which the user did not type.

Where engines do cite, the source mix in Indian AI shopping answers looks nothing like a classic search results page. Counting every appearance across the 4,640 citations in the corpus, youtube.com leads with 534 and reddit.com is second with 191, ahead of nykaa.com (176), and a long tail of small Indian review blogs out-cites most mainstream publishers. Counted instead by how many answers cite a domain at all, reddit.com leads with 106 of 480 and youtube.com follows with 100, so the ranking depends on the basis and we state which one we mean. Our guide on how to get cited by AI covers how to show up on the kinds of sources engines cite in India.

5Which brands gain and lose in Hinglish?

Aggregate overlap statistics are the rigorous part, but the practical question for a brand is simpler: does my visibility change? On the language-sensitive engines, by up to 17.5 points.

Brands that gain or lose visibility in Hinglish: Perplexity
Brand
Change
The Derma Co
+17.5
Taneira
−15
Plum
−13.7
La Roche-Posay
−12.5
Deconstruct
+12.5
Minimalist
+11.3
Vastranand
−10
Garnier
+10
Change in the share of 80 responses per language that mention each brand, for the ten brands that move most. Bars right of the centre line mean the brand is named more often when the question is asked in Hinglish. Percentages are shown with the underlying counts because a single response is 1.25 percentage points.

A pattern worth naming, though we did not test it as a hypothesis and it should be treated as an observation: on Perplexity the brands losing ground in Hinglish skew international or premium, while the gainers skew mass-market Indian direct-to-consumer. The India-origin share of recommended brands rises in the Hinglish arm on Gemini (75.6% to 82%) and Perplexity (63.7% to 66.3%), and stays flat on ChatGPT (74.4% to 74.7%).

Brands that gain or lose visibility in Hinglish: Gemini
Brand
Change
Dot & Key
+11.3
CeraVe
−10
Allen Solly
+10
The Derma Co
−8.8
Zodiac
−8.8
Sparx
+8.8
Change in the share of 80 responses per language that mention each brand, for the ten brands that move most. Bars right of the centre line mean the brand is named more often when the question is asked in Hinglish. Percentages are shown with the underlying counts because a single response is 1.25 percentage points.

Which brand gets named first also becomes less predictable across languages. On Gemini, two same-language reruns agree on the first brand 63.2% of the time, but an English and a Hinglish answer agree only 42.2% of the time. On Perplexity the same figures are 69.5% and 49.8%.

6What should brands selling in India do with this?

Track English and Hinglish separately. If your AI visibility tracking asks English questions only, then on Gemini and Perplexity you are measuring a different brand list from the one Hinglish users see. Our guide on how to measure AI visibility explains how to set up both languages and read them side by side.

Measure each engine on its own. In this study ChatGPT barely moved while Perplexity moved 23.4 points, so an average across engines would hide the change on the engine where it matters most.

Depra tracks Hinglish prompts next to English ones on ChatGPT, Gemini, Perplexity and Google AI Overviews, asked from inside India, and your report shows English and Hinglish visibility side by side. If you are comparing tools, this list of AI visibility tools for D2C brands notes which ones track Hinglish. Plans and limits are on the pricing page.

7How was the study verified?

The full frozen methodology, including the deviations log recording every change made after the design was fixed, is published in the repository. Four verification checks are built into the study, and each of them changed it:

  • Prompt review before analysis. An independent reviewer audited all ten pairs for naturalness, equivalence and purchase intent. Three failed and were rewritten, with affected responses discarded and recollected.
  • Adversarial recomputation. Six independent automated verification runs (AI agents working from raw data, without our analysis code) recomputed the headline numbers with their own code. All six reproduced them. None refuted them.
  • Instrument audit. An audit of 26 brand names that double as ordinary words found three defects in our own tooling, including the brand Simple matching the everyday adjective in 35% of its hits. The verification runs caught two more, both answer-language labels. All five were fixed and the analysis re-run before publication.
  • Stability rerun. Two prompts were collected again 30 to 60 minutes later. The direction holds on all three engines, with drift ratios of 0.91, 1.05 and 1.04, where 1.0 means the later answers matched as closely as reruns inside the window.

One detail cuts against our own headline, so it is worth stating plainly. Twenty-six brand names in our lexicon collide with ordinary English or Hindi words, such as the brand Bata against the Hinglish phrase "bata do". We match those case-sensitively. If we relaxed that rule, the ChatGPT gap would inflate from 2.2 to 13.7 points through Hinglish false positives. Any error in our instrument makes the effect look smaller.

8What are the limitations?

Ten prompts in two categories. These results apply to unbranded skincare and fashion buying questions only. One fieldwork window, so this is a snapshot and engines drift over weeks. One Hinglish register, urban and romanized, so Devanagari Hindi and other Indian languages remain untested. Logged-out default surfaces, so signed-in users may see something different. Citations are the sources an engine displays. They do not prove why it named a brand.

The effect sizes differ by engine in a way that fits differences in how each one handles a Hinglish query, but this design does not separate retrieval from generation and cannot attribute the difference to either. That is open work.

Funding and interest

This study was designed, funded and run by Depra, which sells AI visibility tracking and therefore has a commercial interest in the conclusion that language matters. We cannot remove that interest, so we tried to make it irrelevant. The prompts were frozen before collection, the statistical tests were fixed before analysis, four verification checks are logged in the repository, every correction they forced is recorded in the deviations log, and the entire corpus is public so a skeptic can recompute every number. The finding we would most like to be true, a large ChatGPT effect, is the one we report as fragile.

9How can you reproduce it?

Everything needed to check our arithmetic is in the public repository: all 480 raw responses with their complete payloads, the frozen prompt set, the brand lexicon, the analysis code and the verification logs. The analysis is deterministic and byte-reproducible, so one command confirms a clean recomputation.

git clone https://github.com/ifham001/IndicGEO
cd IndicGEO
node analyze.mjs
shasum -a 256 data/analysis/results.json | cut -d' ' -f1
# 6f9a67de4c11f11b3e157b1f21799fce66161393d24b5a043eb4a1c395de1776

Fresh collection will not reproduce that digest, and it should not: engine answers change over time, so a new run is a new measurement. Replications that disagree with us are welcome as issues or pull requests.

10How to cite this study

Depra Research (2026). IndicGEO Study 01: Same question, different brands. English versus Hinglish buying prompts across ChatGPT, Gemini, and Perplexity in India (Version 1.0.0) [Data set]. https://github.com/ifham001/IndicGEO

Code is MIT licensed, data and text are CC BY 4.0. Machine-readable citation metadata is in the repository's CITATION.cff.

Frequently asked questions

Does ChatGPT recommend different brands in Hinglish?

Barely. ChatGPT's brand list changed 2.2 points more between English and Hinglish than between two reruns of the same question. That result is fragile: the two-sided p is .092, 4 of the 10 prompts moved the other way, and the whole effect sits in the skincare prompts. We describe ChatGPT as approximately language-invariant.

Why does Gemini cite fewer sources in Hinglish?

The study measured the drop but does not establish why. Gemini attached at least one source link to 90.0% of English answers (72 of 80) and 41.3% of Hinglish answers (33 of 80). Its Hinglish answers were also about a third shorter and carried 5.2 citations on average against 9.4. The design cannot separate how Gemini searches from how it writes, so the cause is open work.

Which AI engine changes most with Hinglish prompts?

Perplexity. Its brand list changed 23.4 points more across languages than across reruns (one-sided p = .0010), against 7.8 points on Gemini and 2.2 on ChatGPT. Perplexity was also the most consistent engine on reruns, so the gap stands well clear of its own rerun noise.

How was the study run?

We wrote 10 buying questions (5 skincare, 5 fashion) in English and in Hinglish, none naming a brand. Each version ran 8 times on ChatGPT, Gemini and Perplexity, geolocated to India, inside one 57-minute window on 14 Aug 2026. That gives 480 answers. The full methodology lists every control.

Can I check the data?

Yes. Every raw answer, the prompts, the brand list, the analysis code and the verification logs are public on GitHub. Code is MIT licensed and data is CC BY 4.0. One command, node analyze.mjs, recomputes every number on this page.

Is your brand visible in both languages?

A 3-day trial tracks your brand on ChatGPT, Gemini, Perplexity and Google AI Overviews, in English and Hinglish.

Start 3 days free