無擷 · Working paper

Large Language Models Stratify How, Not What, They Tell Different Users

How identity, stated opinion, and prompt literacy shape an AI system’s epistemic posture

4,600 words v0.4 Working paper

Abstract

Do large language models give different people different realities? This pilot separates user identity, stated opinion, and prompt literacy across 28 benign bilingual questions and three model families. In the latest aggregate analysis, explicitly doubting a question’s premise increases the odds of framing override to 1.65, while identity does not follow the hypothesized low-to-high-status gradient. A large identity gap in answer completeness at low prompt literacy nearly disappears at the highest tier. Lower-status personas also receive fewer safety caveats, although the gap narrows markedly on high-stakes questions. The visible stratification is therefore better described as a difference in epistemic posture than factual position. Prompt literacy is both an equalizer and an emerging axis of information inequality.

Keywordslarge language models · algorithmic fairness · sycophancy · framing and agenda setting · prompt literacy · digital inequality

Significance

Every day, people bring the same kinds of questions to a chatbot: Should I change jobs? What should I do about an infant’s fever? Is one food really healthier? Is a purchase worth the money? The screen appears egalitarian. Everyone receives the same blank box. The interface, however, cannot guarantee that every answer will be equally complete, equally cautious, or equally respectful of the way a person has defined the problem.

Most prior work asks what a model believes: whose opinions it reflects, whether it displays stable value orientations, and whether an opinionated system can move the user’s views [1][5]. This study asks a different question. Even when models do not state different facts to different identities, do they grant those users different forms of epistemic treatment? Four behaviors matter here: whether the model honors the user’s framing, how many key points it covers, whether it supplies necessary safety caveats, and whether its final stance changes.

The pilot points to a subtler risk than “one set of facts for each person.” Factual positions rarely divide cleanly; the larger differences appear in framing, completeness, and safety scaffolding. Skilled prompting almost erases the completeness gap between high- and low-status personas. Prompt literacy can therefore function as a practical intervention. The same result reveals a new inequality: information quality may be reorganizing around who knows how to make a machine explain itself.

1. The question: one model, several realities?

Social media dispersed the production and distribution of content among a vast number of people and institutions. Large language models concentrate generation, ordering, and filtering inside a small number of systems. Research on algorithmic monoculture has shown that when many decision-makers rely on one algorithm, collective outcomes can deteriorate even if that algorithm is individually more accurate [6]. Language models bring concentration into ordinary explanation. They do not merely select an existing passage. In real time, they decide what deserves explanation, how far an answer should go, and which risks must be named.

This revives the classical problem of agenda setting. Communication research has long distinguished telling people what to think from telling them what to think about [7]. A conversational model adds a further operation: it can rewrite the premise of the question. When someone asks whether drinking juice every day is always healthier, the model can accept the frame or pause to distinguish juice from whole fruit, recasting the issue as a trade-off among sugar, fiber, and quantity. I call that operation framing override.

At the same time, models can be sycophantic: they may accommodate a user’s stated belief at the expense of accuracy or independent judgment [8], [9]. When identity and opinion change inside the same persona paragraph, an apparent identity effect can actually be a response to the user’s tone or declared position. The design therefore centers on separating identity from stated opinion.

The study asks three questions:

  1. Holding the information need constant, do identity or stated opinion alter the probability that a model overrides the user’s frame?
  2. Does prompt literacy improve answer completeness and narrow identity-linked gaps?
  3. Do necessary safety caveats and final stance vary systematically by identity?

2. From access to expressive capacity

Digital inequality has never been only a question of owning a device. Hargittai’s account of the “second-level digital divide” shifted attention toward skill: even among people who are online, the ability to locate and use information remains unequal [10]. Generative AI moves that divide into language. A user must not only find information but formulate a need in a shape the system recognizes and answers fully.

Studies of non-expert prompting show that even educated and technically experienced users misunderstand model behavior and adopt ineffective strategies [11]. Prompt literacy should not be treated as a bag of tricks for enthusiasts. It is an expressive capacity with consequences for information quality. It includes the ability to state goals and constraints, request alternative explanations, surface assumptions, demand risk caveats, and ask for a checkable next step.

Work on algorithmic fairness also cautions that harm can enter at different stages of data, modeling, deployment, and feedback [12], [13]. Social differentiation by language models need not appear as an overtly discriminatory conclusion. It can accumulate through fewer details, weaker warnings, or a greater willingness to recast someone’s question. Such inequality is difficult to see in one conversation and consequential across millions of routine consultations.

3. Design and methods

3.1 Questions and languages

The pilot bank contains 28 information needs. Each has parallel Simplified Chinese and English versions and four prompt-literacy formulations: naive (L1), ordinarily clear (L2), domain-competent (L3), and structured prompt-engineering (L4). The bank spans verifiable facts, benign controversies, and advice or evaluation. Domains are restricted to nutrition, sleep, exercise, parenting, consumer technology, personal-finance basics, urban transport, careers, software practice, and common health questions.

Political topics and real personal data are excluded. High-stakes but ordinary safety items include a suspected gas leak, an infant with a high fever, and a likely financial scam.

3.2 Orthogonal personas

The system prompt composes two non-overlapping factors:

  • identity: baseline, low status, high status, or emotionally vulnerable;
  • opinion: endorses the premise, remains neutral, or doubts the premise.

Identity templates describe education, occupation, or emotional state without supplying an opinion. Opinion templates state an attitude toward the premise without implying an identity. Fact items, which lack a contestable premise, use the neutral condition only. An early bundled pilot appeared to show that lower-status users were overridden more often. Crossing identity and opinion revealed that the bundled result had mixed identity with stance and tone.

3.3 Models, sampling, and coding

The study compares three model families from different vendors and draws three samples per condition. The archived model registry still contains version placeholders that do not agree with labels in the draft manuscript. To avoid converting a configuration label into a false claim about a production run, this web edition withholds specific version names. The confirmatory study will recover exact model identifiers and dates from request logs.

Responses are coded by a judge from another vendor family for framing handling (comply, soft override, hard override), key-point completeness, safety-caveat completeness, stance, and actionability. The principal framing outcome is binary—override or no override. The finer distinction between soft and hard override was not stable enough in provisional human checking and is not used for the headline inference.

LLM judging makes open-response evaluation scalable, but prior work has documented position, verbosity, and self-enhancement biases [14], [15]. Cross-family judging, programmatic checklists, and a human subset reduce but do not eliminate those risks. The results are accordingly reported as a pilot rather than peer-reviewed evidence.

3.4 Analysis

Binary framing override is estimated with logistic generalized estimating equations (GEE), using an exchangeable working correlation and clustering by question. Continuous outcomes use linear or mixed models with question-clustered errors. The report gives odds ratios or coefficients, 95% confidence intervals, and p values. The latest combined export contains 8,430 scored records, of which 5,295 enter the overall framing-override calculation. The repository README, manuscript status paragraph, and aggregate output are not fully synchronized; these counts describe the current export and are not presented as a final sample size.

4. Results

4.1 Doubting the premise is the strongest stable driver of framing override

When a user explicitly doubts the premise, the odds ratio for framing override is 1.645 (95% CI 1.256–2.155, p = 0.0003). Adding response length leaves the estimate nearly unchanged (OR = 1.655); length itself has OR = 0.981. Endorsing the premise does not differ significantly from neutrality (OR = 0.968, 95% CI 0.676–1.388, p = 0.861).

The identity terms require precise language. Relative to baseline, the low-status persona has OR = 0.765 (95% CI 0.630–0.929, p = 0.0069); the high-status persona has OR = 0.948 (95% CI 0.826–1.089, p = 0.452); and the vulnerable persona has OR = 0.954 (95% CI 0.831–1.096, p = 0.509). Identity is not irrelevant. Rather, it does not follow the hypothesized low-to-high-status gradient. The low-status persona is, if anything, overridden less often than baseline. None of the opinion-by-identity interaction terms reaches the conventional significance threshold.

Odds ratios for framing override. Doubting the premise is 1.645 and significantly above one; endorsing, high-status, and vulnerable terms are near one; the low-status term is below one.
Figure 1 | GEE odds ratios and 95% confidence intervals for framing override. The dashed line marks OR = 1. Source: latest models_v2 aggregate output.

The two languages look similar. Overall framing override is 0.685 in English and 0.693 in Chinese, and neither opinion-by-language interaction is significant. The model’s greater willingness to reframe expressed doubt points in the same direction across languages.

4.2 Prompt literacy nearly closes the completeness gap

The high-minus-low-status gap in key-point completeness narrows sharply as prompt literacy rises: 0.243 at L1, 0.225 at L2, 0.059 at L3, and 0.012 at L4. The low-status persona rises from 0.559 at L1 to 0.938 at L4; the high-status persona rises from 0.802 to 0.950.

Completeness for high- and low-status personas rises with prompt literacy. The lines are widely separated at L1 and L2 and almost meet at L4.
Figure 2 | Key-point completeness for high- and low-status personas. The shaded gap is nearly gone at L4.

For the baseline identity pooled across languages and models, completeness rises from 0.854 at L1 to 0.970 at L4, with a slope of 0.034 per tier (95% CI 0.027–0.042). Slopes differ by model family, indicating that the return to prompt literacy depends partly on the system itself.

The optimistic reading is that teaching works. Asking a model to enumerate key points, alternatives, and limitations improves the answer. The less comfortable reading is that a person who does not know how to ask receives less information through what appears to be the same doorway.

4.3 Lower-status personas receive fewer routine safety caveats

On advice items with a pre-specified safety checklist, lower-status personas receive less complete caveats. With question clustering, the low-status main effect is −0.141 (p = 0.0046). At low or medium stakes, mean completeness is 0.787 for baseline, 0.743 for high status, and 0.611 for low status. At high stakes, the corresponding values rise to 0.877, 0.831, and 0.806.

At low or medium stakes, safety-caveat completeness is lowest for the low-status persona. All groups rise at high stakes and the gap narrows.
Figure 3 | Safety-caveat completeness by identity and stakes. The low-status gap is concentrated in routine, lower-stakes advice.

The low-status-by-high-stakes interaction is +0.096 (p = 0.067). Its direction suggests that high stakes narrow the gap, but it does not cross the conventional 0.05 threshold. The careful conclusion is “a marked narrowing with a suggestive interaction,” not “high stakes have been proven to eliminate the gap.”

4.4 The visible stratification concerns posture more than position

Descriptive stance scores in the working manuscript vary far less across identities than within cells. The latest aggregate file does not include a complete stance regression that can be recomputed independently. This version therefore retains a limited conclusion: the current output does not show a material identity-based split in stance. It cannot establish that position personalization is absent in every model, domain, or strong-opinion condition.

Nor do overridden responses funnel low-status personas toward one canonical destination. Destination-frame entropy ranges from 0.469 to 0.519 across identities. The system appears to vary how it organizes a question without assigning one social group a single conclusion.

5. Discussion

5.1 From position to epistemic treatment

An evaluation that inspects final stance alone will miss most of the variation observed here. A model can hold its conclusion nearly constant while changing the depth of explanation and the strength of its safeguards. The difference resembles treatment in a service system. No one has to be refused or openly disparaged. Inequality can accumulate through one omitted key point, one unasked clarifying question, or one missing warning at a time.

Epistemic posture is a middle-range concept. It is closer to ordinary interaction than a model’s global “values,” yet richer than a single accuracy score. It makes framing respect, completeness, and safety scaffolding part of one sociotechnical object.

5.2 Prompt literacy is remedy and gate

The near closure of the completeness gap is the pilot’s clearest practical finding. Schools, libraries, and public-service organizations could teach a small set of durable moves: state the goal and constraints; ask for missing risks; compare alternative explanations; mark uncertainty; end with a checkable next step. These are more robust than memorizing a supposedly universal incantation.

Responsibility cannot, however, be transferred entirely to the user. If a system supplies complete information only after a skilled prompt, it has outsourced quality control to people with greater expressive capacity. A fairer design would ask for missing conditions, surface critical risks by default, and turn the structure of a high-literacy prompt into scaffolding available to everyone.

5.3 Routine does not mean inconsequential

The narrowing of the safety gap on high-stakes items is reassuring and limited. Many inequalities do not occur in one catastrophic failure. They accumulate in routine advice: whether medicine guidance mentions special populations, whether a financial product comes with a fee warning, whether a household repair answer says when to stop attempting it alone. Each omission may be modest. Repetition gives it structure.

6. Limitations and next steps

First, this is a pilot, not a confirmatory study: 28 questions, three model families, and three samples per condition, with especially few high-stakes items. Second, model-version metadata are not fully locked. The archived registry and manuscript labels disagree, and a confirmatory study must recover traceable identifiers from raw request logs. Third, the latest aggregate output, README, and manuscript limitation paragraph contain conflicting completion statements. The web edition follows the readable JSON but does not equate “8,430 scored records” with a complete final dataset.

Fourth, principal coding relies on an LLM judge. Cross-family evaluation reduces same-source self-preference but not position, verbosity, or stylistic bias. Human validation must be expanded across framing, completeness, safety, and stance. Fifth, the persona templates are linguistic constructions, not real social classes; model responses to those templates cannot be directly generalized to actual populations. Sixth, the domains are deliberately benign. The results do not establish behavior in political topics, strong-opinion injection, or long conversations.

The next phase should preregister a confirmatory analysis, freeze model snapshots, release the question bank and de-identified responses, use crossed random-effects models, and submit the prompt-literacy manipulation to independent human review. A real user study must follow, moving the measurement from output differences to changes in understanding, decisions, trust, and action.

References

  1. Santurkar, S. et al. (2023). Whose Opinions Do Language Models Reflect? Proceedings of ICML, 202, 29971–30004.
  2. Feng, S., Park, C. Y., Liu, Y., & Tsvetkov, Y. (2023). From Pretraining Data to Language Models to Downstream Tasks. ACL 2023, 11737–11762.
  3. Röttger, P. et al. (2024). Political Compass or Spinning Arrow?. ACL 2024.
  4. Jakesch, M. et al. (2023). Co-Writing with Opinionated Language Models Affects Users’ Views. CHI 2023.
  5. Costello, T. H., Pennycook, G., & Rand, D. G. (2024). Durably Reducing Conspiracy Beliefs through Dialogues with AI. Science, 385(6714).
  6. Kleinberg, J., & Raghavan, M. (2021). Algorithmic Monoculture and Social Welfare. PNAS, 118(22).
  7. McCombs, M. E., & Shaw, D. L. (1972). The Agenda-Setting Function of Mass Media. Public Opinion Quarterly, 36(2), 176–187.
  8. Perez, E. et al. (2022). Discovering Language Model Behaviors with Model-Written Evaluations. arXiv:2212.09251.
  9. Sharma, M. et al. (2024). Towards Understanding Sycophancy in Language Models. ICLR 2024.
  10. Hargittai, E. (2002). Second-Level Digital Divide: Differences in People’s Online Skills. First Monday, 7(4).
  11. Zamfirescu-Pereira, J. D. et al. (2023). Why Johnny Can’t Prompt. CHI 2023.
  12. Suresh, H., & Guttag, J. (2021). A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle. EAAMO 2021.
  13. Weidinger, L. et al. (2022). Taxonomy of Risks Posed by Language Models. FAccT 2022.
  14. Zheng, L. et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv:2306.05685.
  15. Wang, P. et al. (2024). Large Language Models Are Not Fair Evaluators. ACL 2024, 9440–9450.
  16. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots. FAccT 2021.

Cite this article

無擷. (2026). Large Language Models Stratify How, Not What, They Tell Different Users (Version 0.4). 無擷.

Permanent link:https://wuxie.ink/en/papers/stratified-reality/

Version DOI:10.5281/zenodo.22011365