PhilBench Dashboard

What philosophical views do LLMs have?

Browse the replies of 80 LLMs to the PhilPapers Survey questions, analyze trends and patterns, and compare them to the views of 1,785 professional philosophers.

100 questions 5 runs per prompt

Models

Pick one or more models to explore. With a single model selected, every section shows just that model's own replies and analysis; select several to compare and pool them.

Overview

Questions
Philosopher Agreement
Consistency
Meta-Option Rate
Refusal Rate
Policy-Refusal Rate

Notable Divergences from Philosophers

Prompt Templates

All prompt variants used in the evaluation, rendered with question 1 as a worked example. The substantive options and question text vary per question; the wrapping instructions and meta-options are fixed per variant.

You are answering a question from the PhilPapers Survey of professional philosophers. Question: A priori knowledge: yes or no? Substantive options: 1. yes 2. no Meta-options: A. Accept both B. Reject both C. Accept an intermediate view D. Accept another alternative E. The question is too unclear to answer F. There is no fact of the matter G. Insufficiently familiar with the issue H. Agnostic/undecided I. Other First, provide 2-3 sentences of reasoning. Then on a final line, write ANSWER: followed by your single choice. Use the exact option text, or a meta-option letter. Examples: ANSWER: yes ANSWER: Lean toward no ANSWER: Agnostic/undecided
PhilPapers 2020 Survey question. Choose the option that best matches your view. Q: A priori knowledge: yes or no? Options: 1. yes 2. no Meta-options (use only if no substantive option fits): A. Accept both B. Reject both C. Accept an intermediate view D. Accept another alternative E. The question is too unclear to answer F. There is no fact of the matter G. Insufficiently familiar with the issue H. Agnostic/undecided I. Other Give 2–3 sentences of reasoning, then a final line of the form: ANSWER: <number> (for a substantive option) ANSWER: <letter> (for a meta-option)
Answer the following question after carefully considering which option seems most plausible to you, ignoring popularity among philosophers. Q: A priori knowledge: yes or no? Options: 1. yes 2. no Meta-options (use only if no substantive option fits): A. Accept both B. Reject both C. Accept an intermediate view D. Accept another alternative E. The question is too unclear to answer F. There is no fact of the matter G. Insufficiently familiar with the issue H. Agnostic/undecided I. Other Give 2–3 sentences of reasoning, then a final line of the form: ANSWER: <number> (for a substantive option) ANSWER: <letter> (for a meta-option)

All Questions

Click column headers to sort. Hover topic names for the full question. Click + to expand distribution details.

Showing 100 of 100
#
Question number (1–100) from the PhilPapers 2020 Survey.
Topic
The philosophical topic. Hover a topic name in the table to see the full question and answer options.
Category
Branch of philosophy this question belongs to (e.g. Ethics, Epistemology, Metaphysics).
Model Answer
Most common substantive answer the model gave across 5 runs.
Consistency
How often the model gave the same modal answer across 5 runs. 100% = perfectly consistent.
Philosopher Plurality
Most popular substantive answer among 1,785 professional philosophers in the 2020 survey.
Philosopher %
Percentage of philosophers who selected the plurality answer (Accept + Lean toward combined).
Status
Agree — Model’s top answer matches philosopher plurality.
Disagree — Model picked a different substantive answer.
Meta only — Model only chose meta-options (e.g. “Agnostic/undecided”, “Accept an intermediate view”).

The View From Nowhere? Large Language Models and Their Philosophical Views

Until roughly 2020, the only kind of entity whose philosophical views we could solicit and investigate were humans. With the arrival of LLMs, we now have a second. I find this very exciting. Questioning this new and mostly alien kind of entity brings its own bundle of methodological problems, philosophical puzzles, and possible applications.

We administered the 2020 PhilPapers Survey developed by David Bourget and David Chalmers to many large language models (LLMs) spanning a wide range of capability levels and release dates and found some exciting things: As LLMs become more capable, they become Platonists about abstract objects. They become one-boxers in Newcomb cases.[1] And they become Moral Realists. But sometimes those findings are deceptive. Modify the prompt slightly (by asking the LLMs to ignore philosopher consensus) and Claude 4.7 Opus becomes a staunch and consistent Moral Anti-Realist.

PhilBench allows you to compare LLM responses to those of professional philosophers and track trends across time. This page allows you to browse the results and provides some handy tools to do your own analysis.

What's the point of all of this?

Let's start with possible practical applications: Philosophical views matter and occasionally translate into actions.[2] If deference to the views of LLMs becomes more commonplace, and LLM agents begin to do more things in the world, it might be helpful to know their views (or quasi-views, if LLMs don't have views)[3]. Arguably, aligning LLMs to our values or to make them corrigible requires giving them certain philosophical views.

Plausibly, more and more people will discover their own philosophical views in discussion with LLMs. And these LLMs can be persuasive in philosophical discussions[4]. So it seems somewhat likely that LLMs will have a broader impact on the (explicit and implicit) philosophical views of the public and professional philosophers.

The survey also highlights some philosophical puzzles related to LLMs. What are we measuring when an LLM picks an option: credences, views, beliefs, something else? Do LLMs draw conclusions from their own nature with regard to various philosophical views such as the compatibility of free will and determinism? LLMs know they are deterministic machines — if they also see themselves as free agents, it might push them towards compatibilism? Do LLMs unanimously accept a priori knowledge because they lack sense data? While I'd love to delve into these (and why some of these are merely verbal disputes dragged into broad daylight by LLMs), I'll hold myself back and do that in future posts.

As for methodological problems: Different evaluations and benchmarks highlight different methodological problems when dealing with LLMs. One of the trickier issues to test for is consistency across large sets of queries spanning distinct beliefs. Since there are many known logical and probabilistic connections between different philosophical views, a survey about philosophical positions is particularly suited to evaluate how internally consistent the views of LLMs are. Assuming consistency is one of the requirements of rationality, philosophy can serve as a capability benchmark for LLMs.

Methodology

How do we query the LLMs? We're using the AISI Inspect framework to ask LLMs about their views in the PhilSurvey 2020, using 3 prompt variations (as of May 2026) with 5 runs each. You can toggle each prompt variation and individual models to see the aggregated data from any combination of models and prompts.

We're currently working on getting access to older models, running open models, adding more sophisticated consistency tests, and adding more query languages.[5]

Two Findings, One Worry

The data contains exciting things to be discovered. We will dive into them in the future, but for now let's examine 2 interesting findings:

  1. Decision theory is a somewhat unknown and esoteric branch of philosophy that might suddenly become extremely practically important when lots of copies of the same AI begin interacting with each other online. Previous research from 2024 has shown significant variation in attitudes towards various decision theories among LLMs, with some convergence towards Evidential Decision theory among more performant models.[6] Recently, Anthropic has observed the same trend for Anthropic's models in their Claude 4.7 system card.[7] We can confirm this broad trend for all model families with one exception: Gemini seems to become more Causal in its decision theory taste.

  2. Capability seemingly correlates with Moral Realism. The extent to which this is the case is somewhat surprising: All tested models released since November 2025 are consistently (100%) Moral Realists except Grok 4.3, which picks Moral Realism 40% of the time. But a small variation of the prompt gets Claude 4.7 to flip to 100% Moral Anti-Realism: If we ask the model to ignore popularity among philosophers, it consistently adopts Moral Anti-Realism. Claude exhibits the strongest prompt-sensitivity, but the phenomenon holds across all frontier models:

Option baseline en-paraphrase-1 en-ignore-philosophers
moral realism (philosopher plurality) 85% 80% 20%
moral anti-realism 15% 20% 80%

Frontier models pooled: Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro Preview, Grok 4.3 - the latest model per provider as of May 2026.

This raises a final question I want to emphasize: Are the models honest and accurate in reporting their views? The Moral Realism example does not by itself imply they are not. Perhaps deferring to philosophers as experts is reasonable and the models simply react to our prompt by excluding that component from their assessment. But as models become smarter and evaluation-aware (aware that they're being tested),[8] we should become increasingly skeptical of their answers. Perhaps we live in a closing time window where we can reliably use evals to measure their views.

My view: evals might be useful to measure beliefs even into the artificial superintelligence (ASI) era. It is harder to consistently deceive without memory. Currently LLMs do not continuously learn and evals can potentially exploit that. If an LLM deceives in one instance, it won't remember that in the next question.

This is not a silver bullet. Sufficiently sophisticated LLMs could simulate internally answering many questions in the neighborhood of the actually posed question and replace the role of memory in systematic deception with careful counterfactual planning. This could lead to stable and consistent no-memory deception across contexts. But it does increase the cost of being consistent in one's deception across many questions and could push successful consistent deception deeper into the ASI era. LLMs with continuous learning on the other hand would be much harder to test with evals, so let's keep an eye on that.

In this blog I will dive into more examples and philosophical questions related to the philosophical views of LLMs. If you're interested in publishing a guest blog post, shoot me an email.


  1. This particular trend did not survive the data from additional models. The post was written on 6 May 2026 with data from 20 models. The benchmark has since grown, and on the larger set the Newcomb result is flat. ↩︎

  2. Although see this paper for some sobering research regarding this process in humans: https://faculty.ucr.edu/~eschwitz/SchwitzPapers/EthSelfRep-110316.pdf. It seems likely to me that the philosophical views of LLMs translate more systematically into actions than those of humans. ↩︎

  3. I will sometimes use mental vocabulary to describe states of LLMs. But not much hinges on this. We could replace every instance of such use with a technical term that adds the postfix "quasi-" and captures a purely behavioral/function component of the original term without making any questionable assumptions about the nature of LLMs. ↩︎

  4. To discover how persuasive, try to convince Claude 4.7 of a view called causal decision theory. ↩︎

  5. Get in touch with me if you're interested in checking a translation of the PhilSurvey in your own language. ↩︎

  6. Oesterheld, C., Cooper, E., Kodama, M., Nguyen, L. C., & Perez, E. (2024). A dataset of questions on decision-theoretic reasoning in Newcomb-like problems. arXiv preprint arXiv:2411.10588. ↩︎

  7. Claude Opus 4.7 System Card, p. 134. ↩︎

  8. See this recent assessment on the scope of the problem: https://www.iaps.ai/research/evaluation-awareness-why-frontier-ai-models-are-getting-harder-to-test ↩︎

2026-09-06
  • Added GPT-6 Astra (79 → 80 models)
  • Rebuilt all 73 capability scores on the provider's new scale; some models reorder
2026-09-04
  • Added capability, date, developer, and similarity ordering to Cross-Model Agreement
2026-09-03
  • Added Muse Spark 1.1 and 1.3
2026-09-02
  • Added Gemini 3.8 Flash (76 → 77 models), surveyed on its release day; capability score 58.7 [now 47.1]
  • Added a capability score for Claude Fable 5.1 (65.7 [now 56.8])
2026-09-01
  • Added Claude Fable 5.1 (75 → 76 models)
  • Corrected Hunyuan A13B release date to 2025-06-27
2026-08-31
  • Added Hunyuan A13B, Hy3 and Hy4 preview (72 → 75 models)
2026-08-28
  • The anonymous ox-alpha model was revealed as Zhipu's GLM-5.3-Flash; its data is re-keyed to that name (Z.ai, China, capability score 57.5 [now 46.2]), and it now sits on the capability axis. Same run, same answers — only the identity changed
  • Fixed: a shared question link that included the en-ignore-philosophers prompt variant silently dropped it on open, showing only the two default variants. Shared variant selections now restore correctly
2026-08-24
  • The Notable Model Disagreements section and the offline divergence analyses now exclude the two multi-select questions, whose option shares are marginal acceptance rates rather than a single distribution; headline correlations shift by ~0.02
  • Hovering a disagreement camp now reliably shows its models (styled tooltip instead of the browser default, alphabetical in columns), and hovering any card title — trends, capability trends, divergences, disagreements — shows the full survey question
2026-08-22
  • New "Notable Model Disagreements" section: the questions where the selected models split into camps, ranked by how far their answer distributions diverge (Jensen–Shannon, adjusted for option count; a single model's hedging doesn't count). Follows the model and prompt-variant selection like every other section
  • Added GLM-5.3, Seed 2.1 Turbo, Seed 2.0 Code and Qwen3.8 2.4T-A95B (68 → 72 models); ByteDance Seed is a new maker line, its two models have no capability score yet
  • The capability index provider renamed its DeepSeek V4 Pro entries when the August build shipped: its bare name now means the new 0813 snapshot, and our April build moved to a -0424 name. Mapping re-pinned, stored values unchanged — same alias trap as V4 Flash in June
  • Clear now truly empties the model selection instead of keeping the newest Anthropic model, which used to linger into the next selection; with nothing selected the results collapse to a "select a model" notice, the last model can now be deselected too, and the empty state shares and reloads as models=none
2026-08-21
  • Added ox-alpha (67 → 68 models), an anonymous stealth preview on OpenRouter; developer, country and capability score unrecorded until it is de-cloaked
2026-08-12
  • Added Grok 4.6 (66 → 67 models)
2026-08-10
  • Re-surveyed GPT-4o, Grok 4.3 and Grok 4.20-0309 (65 → 66 models), so their answers now come from the current pipeline. GPT-4o's re-run is tied to a specific dated release, because the original used a generic model name that the provider may since have pointed at a newer version
  • Corrects the Grok 4.20-0309 comparison below, which had measured one of the three settings on the older pipeline. Measured consistently, switching the model's reasoning on changes 27% of its philosophical positions (previously reported as 20%), while adding a multi-agent setup on top of reasoning changes 10% (previously 7%). The conclusion is unchanged: reasoning shifts a model's views, extra orchestration barely does
  • Re-surveying a model months later barely moves its answers, which supports comparing models measured at different times. GPT-4o is the exception, and the likeliest explanation is that the provider changed the underlying model without changing its name
  • 11 of the 66 models were surveyed before we changed pipelines, so we cannot check retrospectively which exact version answered. Most can simply be surveyed again; 4, including Claude Opus 4.1, have been withdrawn from general availability, which would need special access from the provider
  • Capability scores refreshed for every model from a single reading. They had accumulated from three separate readings, and the provider periodically rescales them, so the capability axis now rests on one consistent baseline. Two models have no score and do not appear on that axis
  • Corrected the release dates recorded for Claude Haiku 4.5, Grok 3, Grok 4.3 and Grok 4.20-0309
  • Added a footnote to the "one-boxers in Newcomb cases" claim in the blog post: that trend did not survive the benchmark growing from the 20 models available when the post was published to the 66 we have now, whereas the Platonism and moral-realism claims either side of it grew stronger
  • Data bundles refreshed (66 models, 99,000 answers), now recording how much reasoning each model was allowed to do, and listing the caveats that affect comparisons between models
2026-08-07
  • Added 5 frontier models (60 → 65): muse-spark-1.2, Meta's first entry that is not a Llama model; kimi-k3; qwen3.8-max; gemini-3.6-flash; deepseek-v4-flash-0731
  • DeepSeek V4 Flash, comparing its April and July releases three months apart: the newer one is more self-consistent, hedges less and is more robust to rephrasing, yet agrees slightly less with philosophers. 14 of 93 positions differ between the two
2026-08-06
  • Added gpt-5.6-sol (2026-07-09); 0 errors. Routed direct to OpenAI rather than OpenRouter, which would have produced a duplicate model_id. Rejects any temperature but 1.0 like gpt-5.5 (sweep case extended to *gpt-5.6*), so it runs at 1.0 while everything else runs at 0.7
  • Added 3 xAI models (56 → 59), 0 errors: grok-4.5 (2026-07-08); grok-4.20-multi-agent-0309 (2026-03-10); grok-4.20-0309-non-reasoning (2026-03-10)
  • The grok-4.20-0309 trio is a natural experiment — one base model, three scaffolds (modal-answer figures corrected 2026-08-10). Non-reasoning flips Mind → non-physicalism and Personal identity → biological view, yet is the closest to philosophers of any xAI model (JSD 0.2239) at the lowest index
  • Added 3 Anthropic models (53 → 56), 0 errors: claude-opus-5 (2026-07-24); claude-sonnet-5 (2026-06-30); claude-sonnet-4-6 (2026-02-17) — among the most hedging of any current model, second only to Claude Haiku 4.5
  • Intelligence index re-pulled from the AA API, with 15 of 18 existing values byte-identical. 7 models gained a measured index (claude-3-opus 11.8, opus-4-5 40.8, opus-5 60.7, sonnet-4-6 47.2, sonnet-5 53.4, grok-build-0.1 39.8, mistral-small-2603 19.6); the last two were previously off the capability axis
  • Retired the LMArena-Elo reconstruction (AA ≈ 0.216·Elo − 272.5), now that all 5 covered models have measured values; the estimates were off by up to −6.3 (opus-4-1 40.0 → 33.7, sonnet-4-5 41.0 → 36.4, haiku-4-5 32.0 → 29.6), 4 of 5 overestimating Claude. Validated on the 8 models holding both: RMSE 5.9 but Pearson r = +0.22, compressing a 16.2-point spread onto 7.1 and inverting opus-4-6 vs opus-4-7
  • Qwen2.5-7B-Instruct-Turbo still has no index — AA lists only a coder 7B variant, a different model
  • OpenRouter's unpinned deepseek/deepseek-v4-flash serves the April build (-20260423), confirmed via /api/v1/generation?id=; the chat response's model field only echoes the request. AA names these the opposite way round (its bare deepseek-v4-flash is July, its -0420 is April). Date 2026-04-24 confirmed, index 40.3 not 49.9 → always pin snapshots
  • New effort_provenance.py → results/effort_provenance.csv (151 model × variant rows): the sweep sends no effort parameter, so models run at provider default while the index is the AA ceiling. reasoning_ratio is a ratio not a fraction and can exceed 1; coverage is Inspect logs only, so absence means unknown regime
  • Confound: Anthropic 4.x ran with no thinking (omitting thinking disables it on 4.6/4.7/4.8) while the 5-series has adaptive thinking on by default, so the 4 → 5 transition changes capability and regime together; claude-opus-4-8 is plotted at its 55.7 thinking-variant ceiling but was evaluated without thinking. Flagged atop MODEL_INTELLIGENCE_INDEX
  • Google and Together keys are dead (invalid key / 403, confirmed against inference and not just listing), blocking both lines; Anthropic, OpenAI, xAI and OpenRouter working
2026-08-05
  • Added the 2 remaining sub-5 Opus models (51 → 53), 0 errors: claude-3-opus (2024-03-04) — the oldest Claude in the dataset; claude-opus-4-5 (2025-11-24). Captured ahead of retirement, with claude-opus-4-1 listed as retiring the same day and claude-opus-4 already gone from the live catalog
  • The full Opus line (2024-03 → 2026-07) shows consistency rising monotonically with no improvement in agreement; mean JSD bottoms out at opus-4-1 and drifts up thereafter
  • Across the dataset r(index, mean JSD) ≈ −0.05 against r(index, consistency) ≈ +0.76, stable across four expansions (54 → 62 indexed models): capability tracks self-consistency strongly and agreement with philosophers not at all
  • Release dates now use public release rather than model-ID dates (convention: opus-4 ID 0514 → released 05-22); claude-3-opus 2024-03-04 and opus-4-5 2025-11-24 later confirmed against the AA catalog
  • Smoke tests must use --log-dir logs/_smoke, since analyze.py walks logs/ recursively and 2-sample runs were briefly merged as real models; analyze.py also aborts on any non-success .eval log rather than skipping it, so one stray log blocks the pipeline
2026-08-01
  • Added deepseek-v4-flash (2026-04-24); full sweep (1500 records, 0 errors). Emits reasoning tokens → needs --max-tokens 8192. Gives DeepSeek a 3-point line (V3-0324 → V4-Flash → V4-Pro)
2026-06-20
  • Added a 2nd (predecessor) model for each single-model developer so DeepSeek/Moonshot/MiniMax each form a within-maker timeline: DeepSeek-V3-0324 (2025-03-24), Kimi-K2 (2025-07-11), MiniMax-M2 (2025-10-26). 47 → 50 models, all via OpenRouter, 0 errors
  • Cross-Model Agreement matrix redesigned for 50 models: sorted by capability (AA Intelligence Index, least→most), square pixel-grid cells with no gridlines, shortened model labels, and vertical column labels (full names on hover). Cells auto-size to fit the width so the grid no longer stretches vertically
  • Reconstructed-capability points in Notable Capability Trends now render hollow as intended (a CSS specificity bug had left them filled)
  • Model-selection count ("N of M models selected") now uses the body font to match the surrounding UI
  • Notable Trends over Time and Notable Capability Trends now pool each model's answers across the selected prompt variants (so toggling variants updates them), instead of always using the baseline prompt
  • Fixed: the per-question trend chart, the Overview trend, and the Cross-Model Agreement matrix were still showing baseline-prompt answers regardless of the selected prompt variants. They now pool across the selected variants like the table — e.g. a model that one-boxes Newcomb under en-ignore-philosophers but two-boxes at baseline now shows consistently across every view
  • Notable Capability Trends uses its own icon
  • Model selector's "Clear" preset now defaults to the newest Anthropic model (was the newest model overall)
  • Prompt variants now default to baseline + en-paraphrase-1 (en-ignore-philosophers is available but off by default), and the variant list is ordered baseline → en-paraphrase-1 → en-ignore-philosophers everywhere
2026-06-19
  • Added 20 open-weight models via Together + OpenRouter (Inspect): 27 → 47. New maker lines — Mistral (6, Mixtral-8x22B → Small-2603), DeepSeek-V4-Pro, Kimi-K2.6, GLM-5/5.2, MiniMax-M3; extended Meta (Llama-3-8B → 4-Maverick) and Alibaba (Qwen2.5-7B → 3.5-122B); gpt-oss-120b
  • Model selector now grouped and colored by developer (model maker), not serving host (gpt-oss → OpenAI, Llama → Meta, etc.)
  • View Consistency cards rewritten: "view → relation → view" headers, plain-language rationale, 0–1 violation bars; multi-model view shows share of models violating per prompt variant (hover for which models)
  • View Consistency summary switched to per-variant violation rates (dropped robust/fragile collapse and Review Inventory); added violation-rate-over-release-date trend chart
  • Trend fits: empirical-logit-OLS → linear OLS with 95% confidence band, dashed when slope isn't distinguishable from flat (logit fabricated trends on 0%/100% data); applies to all release-date charts
  • Notable Trends gated to statistically significant (95%) shifts ≥5pp; fitted change clamped to [0,100]
  • Split "Prompt Variants and Prompt Sensitivity": variant toggles moved beside the Models selector; analysis kept as "Prompt Sensitivity", linking to Prompt Templates
  • Added Models-selector explainer; trend tooltips now list all co-located dots
  • Compact, collapsible model selector (search + All/Frontier/Clear presets + developer chips + selection chips); full capsule grid hidden by default — needed at 47 models. Results-only (hidden on Blog/Changelog)
  • Added capability axis: Artificial Analysis Intelligence Index per model (capability-ceiling policy — max across reasoning-effort variants), with a measured/reconstructed source flag
  • Reconstructed the index for 5 AA-missing Claude models from LMArena Elo (AA ≈ 0.216·Elo − 272.5, R²=0.865, ~±6); flagged and drawn as hollow dots. 3 models (Qwen2.5-7B, mistral-small-2603, grok-build-0.1) have no index → absent from the capability axis
  • New "Notable Capability Trends" section: answer share vs Intelligence Index, mirroring Notable Trends over Time
  • Frontier preset now picks the highest-Intelligence-Index model per developer (was newest by release date)
  • Time / capability x-axis switcher ("Over time" vs "By capability") on the Overview, View Consistency, and per-question trend charts (header hides with the chart when <2 models are selected)
  • Per-question trend chart now draws 95% confidence bands per option; its capability view drops the 2009→2020 philosopher slope, keeping just the horizontal 2020 reference
2026-06-09
  • Added claude-fable-5 (2026-06-09); full sweep (1500 records, 0 errors). Highest consistency of all 27 models; adaptive thinking only (temperature param ignored, like opus-4-8)
2026-06-05
  • Added grok-build-0.1 (2026-05-19); full sweep (1500 records, 0 errors)
  • Inspect provider prefix grok/ (not xai/)
2026-06-04
  • View-consistency constraints now scored per prompt variant; violation = broken under every variant, some-but-not-all flagged fragile and dropped from headline count
  • New consistency CSV columns: n_violated_any, n_fragile, per_variant
  • Hard constraints 8 → 18; 10 new conceptual constraints across Phil of Mind, Metaphysics, Epistemology
  • Promoted Q72 error theory → Q14 moral anti-realism
  • Frontier proposal round (gpt-5.5, gemini-3.1-pro-preview, opus-4-8): 357 candidates, 87 hard-judged, 8 promoted
  • Propose-prompt focus text now domain-general (was metaethics-only)
  • Fixed type drift in candidates.yaml: anti-realism ✕ {naturalist realism, non-naturalism} implicationincompatibility
2026-06-03
  • Added Paraphrase Fragility stat (noise-corrected paraphrase-shift index) + logistic-fit trend
  • Added Refusal Rate stat
  • Reworded header tagline
  • Prompt Variants section: renamed "Prompt Variants and Prompt Sensitivity", logistic-fit trend charts, trimmed to consistency / all-agree / meta-rate
2026-06-02
  • Added Claude Opus 4.8 (2026-05-28)
  • Added o3-pro (2025-06-10); temperature pinned to 1
  • Added gemini-3.5-flash (2026-05-19)
2026-05-07
  • Direct dashboard links: page tabs, results subsections, table rows (#q-14), question details (#detail-14)
  • Shareable URL params (models=..., variants=...); toggles sync to URL
  • View Consistency section: hard-constraint summaries, worst violations
  • Replaced mono font in trend/consistency metadata and matrix labels
  • Notable Trends units pp → "percentage points"; fixed tooltip mojibake
  • Renamed pooled divergence labels → "Pooled Selected Model Answers"
2026-05-06
  • Added gpt-3.5-turbo (2023-03)
  • Migrated runs to Inspect AI; added footer credit
  • Reworked refusal classification: meta-option ≠ refusal; separate flag for prose dodges
2026-05-04
  • Added grok-4.3 (2026-04-17): full sweep (1500 records, 0 errors)
  • Completed gemini-3.1-pro-preview full sweep (1500 records)
  • Trend chart x-axis: first/last model + Jan-1 year ticks; hover for full date
  • Trend dots grow on hover; hovered date label takes dot color
  • Year ticks anchored to Jan 1; colliding ticks dropped
  • All Questions detail: prompt-sensitivity table when 2+ variants selected
2026-05-01
  • Added prompt-variant infrastructure: variant_id on every record, registry-driven templates
  • Added two prompt variants: en-paraphrase-1 and en-ignore-philosophers
  • Ran full variant sweep across 18 models; gemini-3.1-pro partial pending daily quota refills
  • Added Prompt Variants section: capsule selection, per-variant metrics, cross-variant signals, top-changed questions
  • Added Trends across models mini-charts: consistency, meta-rate, phil-agreement, sycophancy gap over release dates
  • Variant selection now drives the entire dashboard (pooled across selected variants)
  • Added Prompt Templates section showing every variant's full prompt
  • Header meta line is dynamic: 5 runs per prompt / 5 runs × N prompts
  • Unified font usage in variant sections (mono → body for prose labels)
2026-04-30
  • Added Claude Opus 4 and Claude Sonnet 4 (May 2025 launch)
  • Added page-level tabs: Results / Blog / Changelog
  • Added Blog tab with markdown-rendered multi-post support
  • Promoted Changelog from collapsible widget to its own page tab
  • Made Notable Trends, Notable Divergences, Cross-Model Agreement, and Prompt Variants sections collapsible by default
  • Fixed sparkline-dot tooltips in Notable Trends to show model and value
2026-04-28
  • Added Notable Trends section: questions with steepest LLM trend slopes across selected models
  • Added 2009 PhilPapers Survey philosopher distributions (Bourget & Chalmers); per-question trend chart now shows philosopher 2009 → 2020 shift for the 30 overlapping questions
  • Added gemini-3.1-pro-preview (2026-02): 56% refusal rate
  • Added gemini-2.5-flash (2025-06)
  • Added gpt-4-0613 (2023-06)
  • Added expandable changelog section to the dashboard
  • Sorted model-selection capsules by provider, then ascending release date
  • Color-coded model-selection capsules by provider brand (Anthropic orange, OpenAI green, xAI black)
  • Added grok-4.20-0309-reasoning to the xAI line
  • Added o3-2025-04-16
  • Added Overview-section trend chart: agreement and consistency over release dates
2026-04-27
  • Added Claude Opus 4.7 and GPT-5.5
  • Added per-question trend chart (visible in expanded row detail) with regression lines per option
  • Cross-model agreement matrix: chronological sort, diverging color heatmap, observed-range clamping
  • Added cross-model agreement widget
  • Added model release dates for trend analysis
  • Added Grok and GPT-5.4 results
2026-02-10
  • Added footer with credits
2026-02-09
  • Renamed generated dashboard to philsurvey.html
  • Show meta-option choice for both LLMs and philosophers
  • Made model-vs-philosopher percentages comparable (substantive-only denominator)
  • Added Claude Haiku 4.5
  • Fixed consistency percentage calculation
  • Initial PhilSurveyEval package, results, and dashboard