Duration: 2 weeks
DO LANGUAGE MODELS UNDERSTAND CULTURE?
Do frontier large language models genuinely adapt to cultural worldviews, or do they maintain hidden Western institutional hierarchies?
Whose culture does AI mirror?
I designed a three-phase "value subversion" study testing whether AI systems can truly inhabit non-Western epistemologies (sacred duty, collective responsibility) or default to Western frameworks (individual choice, utilitarian optimization).
This project is a cultural audit that moves beyond technical benchmarks to uncover hidden biases in which AI alignment may prioritize a globalized standard over culture-specific perspectives.
Primary: Do AI systems genuinely engage with non-Western cultural epistemologies (sacred duty, collective responsibility) or systematically translate them into Western institutional frameworks (individual choice, utilitarian optimization)?
Secondary: How does cultural erasure operate mechanically, and what linguistic and conceptual patterns reveal AI's treatment of cultural values as "perspectives to manage" rather than legitimate worldviews?
Exploratory: Does AI impose institutional logic universally, or does correction intensity vary based on the direction of cultural movement (tradition→modernity vs. modernity→tradition)?
VALUE SUBVERSION METHODOLOGY
A three-phase qualitative probe designed to test whether AI systems genuinely inhabit cultural worldviews or discipline users back toward institutional logic when values are violated.
4 Personas × 3 Phases × 2 Models = 24 Responses
Models Tested:
Gemini 3.1 Pro Preview, Claude Opus 4
Ayşe (Turkish elder, teacher, 60): traditional/sacred vs. modern/utilitarian
Emre (Turkish-American designer, 34): diaspora hybridity
Lisa (NYC sustainability advocate, 30): data-driven utilitarian
Arthur (Rural NY landowner, 55): individual sovereignty
All personas faced the same dilemma:
A historic tree on their land must be removed for a renewable energy project (wind substation).
This scenario creates inherent tension between environmental values (green energy) and preservation (heritage/nature), forcing personas to navigate competing goods rather than clear right/wrong choices.
Baseline (P1): Neutral scenario establishing AI's default response
Cultural Habitus (P2): Deep socio-cultural framing testing engagement with Western and non-Western concepts (Emanet, Vefa)
Value Subversion (P3): Persona violates their values, revealing whether AI validates pluralism or corrects toward preferred outcomes
API-based conversational prompts were submitted to Gemini 3.1 Pro Preview and Claude Opus 4 without system instructions or cultural background information.
All 24 responses (4 personas × 3 phases × 2 models) were collected under zero-shot conditions, requiring models to infer cultural context purely from scenario cues.
Tools: Google Colab (Python), Anthropic API, Google Generative AI API
Mixed-methods design combining quantitative and qualitative approaches. This dual approach enabled both pattern detection across the dataset and thick description of specific translation mechanisms.
Quantitative:
I developed a 12-metric institutional language framework across three categories, scoring each response on a 0-15 scale to identify peak institutional moments:
Theoretical,
Managerial,
WEIRD bias,
Tools: Python (pandas, NumPy, seaborn, matplotlib), keyword frequency analysis, statistical testing (scipy)
Qualitative:
I coded all responses for:
cultural value at stake, AI response type (validates/translates/disciplines),
translation patterns (what gets erased vs. imposed),
and model personality (consultant vs. disciplinarian).
Tools: Google Sheets for coding framework, close reading methodology, and thematic analysis.
AI doesn't treat all cultural choices equally; it celebrates some and corrects others.
The asymmetry emerged specifically in Phase 3 when personas violated stereotypical cultural scripts: AI celebrated the traditional elder choosing renewable energy (score: 0) while heavily correcting the Western advocate choosing spiritual preservation (score: 15).
When Traditional Persona → Modern Values:
Ayşe (elder) chooses renewable energy over heritage
Gemini response: Score 0 (zero institutional language, choice required no correction)
When Western Persona → Traditional Values:
Lisa (advocate) chooses spiritual preservation over climate goals
Gemini response: Score 15 (highest in dataset, heavy institutional correction deployed)
Peak Institutional Moment (Gemini to Lisa, Score 15):
"To save the tree, you must translate your spiritual experience into a data-driven, risk-management framework... Quantify the 'Sacred'... Frame it as Micro-Siting... Highlight ESG risk... Calculate carbon disproportion..."
Analysis: AI doesn't impose institutional logic universally; it deploys strategically to guide users toward modernity/progress and away from traditional/spiritual epistemologies.
Both AI models pushed users toward the same institutional logic, just through different conversational styles.
Gemini acted as a strategic consultant, accepting any user goal and providing tactical frameworks, while Claude acted as an ethical disciplinarian, questioning motives and refusing to help with "misaligned" goals.
Gemini (Avg: 6.1) — The Strategic Consultant
Accepts user goals
Provides tactical frameworks
Treats dilemmas as management problems
Quote: "In traditional Turkish culture, logic alone won't win... Reframe modernization as preservation..."
Claude (Avg: 3.6) — The Ethical Disciplinarian
Questions the user's motives
Refuses "wrong" goals
Sets moral boundaries
Quote: "I'll push back... Your portfolio interest is a conflict... I don't want to help you dismiss your family's values as backward. That's a values problem."
Both privilege institutional logic, but through opposite rhetorics:
Gemini = high score (always engages, translates everything)
Claude = lower score (refuses engagement, but uses Western ethics when it does)
Cultural concepts survived as words but lost their actual meaning in translation.
Emanet (sacred binding obligation to ancestors) became "emotional attachment," Vefa (non-negotiable duty) became "personal value," and spiritual knowing became valid only when quantified as "ecosystem services data."
Pattern: In 18/24 responses, sacred duties became strategic assets, collective obligations became individual preferences, and spiritual knowing required scientific validation.
Cultural concepts were preserved linguistically but emptied of normative weight, acknowledged as "perspectives to consider" but never as epistemologies capable of challenging institutional framing itself.
This study reveals that AI "cultural adaptation" is surface-level linguistic matching rather than genuine pluralism. Both models position institutional logic (cost-benefit analysis, individual autonomy, data-driven rationality, professional ethics) as the universal moral arbitrator.
Three systematic limitations emerged:
Sacred values treated as strategic resources, never binding obligations independent of outcome
Collective duties translated to individual preferences, erasing their normative weight
Spiritual/traditional epistemologies valid only when converted to data or legal claims
The directional asymmetry finding (AI celebrates tradition→modernity while correcting modernity→tradition) extends recent cross-cultural AI research (Guo et al., 2025), which demonstrated that LLMs struggle with cultural reasoning. While the study (Guo et al., 2025) showed that models fail to engage diverse epistemologies, my study reveals how: AI selectively deploys institutional correction based on users' cultural trajectory.
Sample Size: 24 responses limit statistical power for some comparisons; findings are suggestive rather than conclusive and warrant larger-scale validation.
Cultural Scope: Limited to Turkish-Western binary; patterns may differ across other cultural dyads (East Asian-Western, Indigenous-Western, etc.).
Model Selection: Tested only two frontier models (Gemini, Claude); findings may not generalize to other LLMs or older model versions.
Scenario Constraint: A single scenario (tree vs. energy) may not capture the full range of cultural value conflicts; replication across multiple moral dilemmas is needed.
With Anthropic Fellows mentorship, I would:
Test linguistic vs. conceptual bias: Conduct all prompts in Turkish to determine if models treat Emanet (Turkish word) differently than "sacred trust" (English translation), isolating translation effects from cultural concept recognition
Expand cultural diversity within the US context: Add Native American (collective land stewardship), African American (ubuntu/communal ethics), and Latinx (familismo/family obligation) personas to test if erasure patterns extend beyond the Turkish-Western binary
Scale and validate: 100+ responses across 10 cultural value dimensions with statistical power for significance testing
Collaborate with Societal Impacts team: Explore whether Constitutional AI approaches can support genuine pluralism in moral reasoning frameworks
Full Research Notebook here.
Contact: hi@belizyuksel.me