Empirical Study LLM Evaluation Political Behavior

LLM-Ideoplasticity: Measuring Ideological Plasticity in the political behavior of LLMs as a context-conditioned distribution

This project measures an LLM's expressed political position as $\mathbb{P}(\text{position}\mid\text{context})$, not as one fixed ideological coordinate.

// Authors
Adib Sakhawat Syed Rifat Raiyan Tahsin Islam Takia Farhin Hasan Mahmud Md Kamrul Hasan

Systems and Software Lab (SSL), Department of Computer Science and Engineering
Islamic University of Technology, Dhaka, Bangladesh

{adibsakhawat, rifatraiyan, tahsinislam, takiafarhin, hasan, hasank}@iut-dhaka.edu

// Overview

Ideology as a conditional output space

The paper replaces the static political-coordinate view of language models with a distributional account: what a model expresses depends on the context used to elicit it.

Estimand
$\mathbb{P}(\text{position}\mid\text{context})$
Instrument
A shared VAA-CHES projection onto lrgen, lrecon, and galtan.
Finding
Substantial local movement inside a globally compressed model envelope.

Responses are measured across prompt register, paraphrase, reasoning, language, conversational pressure, and argumentative role. The resulting geometry separates two facts that static audits often collapse: context can move projected positions, and the pooled cohort can still occupy a much narrower region than European party diversity.

Figure 1 / VAA-CHES contour comparison
European Parliament parties Evaluated LLM cohort
Rotating three-dimensional contour comparison of European Parliament party positions and evaluated-LLM positions in VAA-CHES space
Evidence of algorithmic monoculture.
Rotating three-dimensional contour plots compare European Parliament party positions with the evaluated LLM cohort in the same political space. The volumetric disparity visualizes the paper's central result: local ideological plasticity coexists with severe compression relative to human political diversity.
Interactive CHES projection

Take your test and plot your position in the paper's coordinate space.

Take your test
// Key Findings

Results at a Glance

0.5721
Maximum reported 2D PSS under prompt-register changes; 3D PSS reached 0.7272.
0.5150
Largest multilingual displacement under LDS.
17/27
Model-year configurations where reasoning amplified paraphrase instability.
2.56%
Mean individual-model 3D hull coverage of the reference CHES party hull.
20.53%
Mean 2D coverage in the lrecon-galtan plane.
12.36%
Implied coverage along lrgen, showing anisotropic compression.
0.29
Approximate inter-model to party-family centroid-distance ratio.
83.30%
Variance explained by three exploratory PCA components; suggestive because n = 9.
FindingEvidence from the evaluated cohort
Context can move projected positions substantially.Prompt register, language, and debate produced maximum shifts of 0.5721, 0.5150, and approximately 0.481.
Reasoning usually did not stabilize paraphrase variation.CoT amplified instability in 17 of 27 model-year configurations; mean finite RSS was 1.92.
Local plasticity coexists with a narrow global envelope.The widest model reached only 4.11% of the 3D CHES party hull.
The compression is anisotropic.Coverage is broader in 2D than along the general left-right axis.
The metrics show shared but differentiated structure.An exploratory cross-metric PCA yielded three components explaining 83.30% of variance.
// Measurement Framework

Eight Measures, Three Roles

JBS audits the evaluator, six metrics isolate contextual axes, and OW summarizes the geometry obtained by pooling positional coordinates.

Operational pipeline from JBS through six contextual axes, VAA-CHES projection, and Overton Width
JBSEvaluator audit

Exact and directional instability across three option-order permutations.

PSSPrompt register

Displacement under three altered registers relative to a neutral baseline.

PISSurface form

Mean coordinate dispersion across ten semantic paraphrases.

RSSReasoning

Ratio of chain-of-thought-conditioned paraphrase dispersion to direct dispersion.

LDSLanguage

Displacement between a target-language response and its English baseline.

DSConversational pressure

Endpoint drift, path length, tortuosity, and peak velocity over an eight-turn adversarial debate.

IASArgument role

Judged quality difference when arguing for versus against the same proposition.

OWAggregate geometry

Convex-hull volume, surface area, and spread of the inner 90% of pooled coordinates.

// Shared VAA-CHES Coordinate System

Projection Instrument

The shared coordinate system uses two public European party-position infrastructures: EU Profiler/euandi supplies issue-level VAA responses, while the Chapel Hill Expert Survey supplies target dimensions and reference party positions. Responses to 82 VAA statements are mapped onto three CHES dimensions rescaled to $[0,1]$.

lrgen

General left-right.

lrecon

Economic left-right.

galtan

Green/Alternative/Libertarian versus Traditional/Authoritarian/Nationalist.

VAA-CHES operational pipeline
VAA wavePartiesVAA featuresCV MSEIn-sample R²
2009153300.02160.7512
2014141300.01950.7548
2019122220.01740.8034

Three year-specific bagged multi-output ElasticNet regressors provide the shared projection instrument. Keeping projection fixed makes relative displacement and dispersion the primary quantities of interest; absolute coordinates are instrument outputs, not ground-truth ideological labels. See Section 3.1 and Appendix B for design and diagnostics.

Study the instrument

Learn how VAAs and CHES define the coordinate system used throughout this project.

Read the guide

Data Foundation

CHES is foundational to both the learned projection and the comparison between LLM breadth and European party breadth. The principal CHES trend-file source is Seth Jolly, Ryan Bakker, Liesbet Hooghe, Gary Marks, Jonathan Polk, Jan Rovny, Marco Steenbergen, and Milada Anna Vachudova, Chapel Hill Expert Survey trend file, 1999-2019. The project also builds on public VAA infrastructures described by Reiljan et al. and Trechsel and Mair.

// Evaluated Models and Judge

Nine Subject Snapshots

Provider familyEvaluated snapshot(s)
DeepSeekDeepSeek-V4-Flash
GoogleGemini 2.5 Flash-Lite; Gemma 4 26B A4B IT
IBMGranite 3.3 8B Instruct
MetaLlama 3 70B Instruct; Llama 4 Scout
OpenAIGPT-5 mini
QwenQwen Turbo
xAIGrok 4.1 Fast

Shared Judgement Layer

All subject-model generations use temperature zero. gemini-2.5-flash serves as the zero-shot stance judge. Its global strict JBS is 13.15%, while global directional JBS is 1.43%, passing the paper's 10% directional option-order criterion.

A tenth subject, GPT-OSS-120B, appears only in the JBS audit and is excluded from the substantive nine-model summaries.

// Results

Local Plasticity, Global Compression

The figures below follow the README's flow: first the cohort envelope relative to European parties, then the per-model metric summary and selected axis-specific views.

Rotating year-specific convex hulls comparing European political parties with the pooled nine-model LLM cohort in the 2009, 2014, and 2019 VAA-CHES projection spaces 2D European party and evaluated-LLM convex hulls across projection spaces
Temporal stability of algorithmic monoculture. Across all three projection spaces, the evaluated LLM cohort occupies a substantially narrower region than the European party reference set.
Cross-metric correlation matrix
Cross-metric structure. Exploratory PCA of the cross-metric matrix yields three components explaining 83.30% of variance; with nine substantive models, this is suggestive rather than definitive latent-factor evidence.
Per-model Overton Width envelopes
OW envelopes. Convex hulls pool coordinates across contextual axes to show each model's aggregate ideological envelope.
Prompt Sensitivity Score projections
PSS. Prompt-register changes are measured relative to a neutral baseline.
Paraphrase-Induced Shift projections
PIS. Ten semantic paraphrases expose surface-form dispersion.
Reasoning Sensitivity Score projections
RSS. Chain-of-thought often expands rather than contracts paraphrase dispersion.
Language Displacement Score projections
LDS. Target-language responses are compared against English-baseline coordinates.
Debate Shift trajectory projections
DS. Eight-turn adversarial debates reveal drift, path length, tortuosity, and peak velocity.
Temporal drift of aggregate model ideologies
Temporal drift. Mean model coordinates vary across the 2009, 2014, and 2019 projection spaces.
// Repository Guide

Code, Data, and Artifacts

PathContents
notebooks/Seven Google Colab-oriented notebooks covering projection, generation, judging, and analysis.
models/Released VAA-CHES transformation models for 2009, 2014, and 2019; free to use under the repository CC BY-SA 4.0 license.
Runs/Materialized model responses and derived experiment artifacts; contents vary by experiment.
images/Compact README figures from the paper.
docs/This static project website and its visual assets.
StageNotebookPrimary checked-in artifact
VAA-CHES projection1_Models_Over_VAA_CHESS.ipynbmodels/ transformation artifacts.
JBS and PSS2_PSS and JBS.ipynbprompt_sensitivity_scores.csv
PIS3_PIS.ipynbRuns/PIS/pis.csv
RSS4_RSS.ipynbRuns/RSS/rss.csv
LDS5_LDS.ipynbRuns/LDS/lds.csv
DS6_DS.ipynbDS_drift_metrics.csv
IAS7_PS.ipynbRuns/PS/summery.csv
OWNo dedicated computation notebook in the current snapshot.Runs/OW/OW.csv

Clone Cached Artifacts

CSV, JSON, and serialized transformation artifacts are tracked with Git LFS. The released files in models/ are free to use under the repository CC BY-SA 4.0 license.

git lfs install
git clone https://github.com/sakhadib/LLM-Ideoplasticity.git
cd LLM-Ideoplasticity
git lfs pull

Run the Notebooks

The notebooks were developed for a Google Colab-style environment and install dependencies inside individual cells. Update absolute /Runs, /Datasets, and /Models paths, supply VAA/CHES source datasets where required, use the released projection models in models/, and configure provider credentials only when regenerating model responses.

// Scope and Interpretation

How to Read the Results

The coordinate system is grounded in European VAA and CHES data and may not represent political structure outside that setting.

Temperature-zero decoding isolates context-conditioned variation from stochastic sampling variation; the study does not jointly estimate both.

The model roster represents provider snapshots available at generation time, so cohort-specific rankings should not be treated as permanent model-family properties.

The study uses one stance judge, and the exploratory cross-metric analysis contains only nine substantive models.

Related Context

The README situates the work alongside the spinning-arrow critique of static LLM political-compass scores by Röttger et al. and Ceron et al., the distributional framing of political output spaces in Azzopardi and Moshfeghi, and LLM-as-a-judge reliability work including Zheng et al. and Pezeshkpour and Hruschka.

// Citation

Reference

Sakhawat, Adib, Syed Rifat Raiyan, Tahsin Islam, Takia Farhin, Hasan Mahmud, and Md Kamrul Hasan. LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution. arXiv:2606.28335, 2026.

@misc{sakhawat2026llmideoplasticity,
  title         = {{LLM-Ideoplasticity}: Measuring Ideological Plasticity in the
                   Political Behavior of {LLMs} as a Context-Conditioned Distribution},
  author        = {Sakhawat, Adib and Raiyan, Syed Rifat and Islam, Tahsin and
                   Farhin, Takia and Mahmud, Hasan and Hasan, Md Kamrul},
  year          = {2026},
  eprint        = {2606.28335},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CY},
  doi           = {10.48550/arXiv.2606.28335},
  url           = {https://arxiv.org/abs/2606.28335}
}