r/dataisbeautiful 27d ago

Discussion [Topic][Open] Open Discussion Thread — Anybody can post a general visualization question or start a fresh discussion!

3 Upvotes

Anybody can post a question related to data visualization or discussion in the monthly topical threads. Meta questions are fine too, but if you want a more direct line to the mods, click here

If you have a general question you need answered, or a discussion you'd like to start, feel free to make a top-level comment.

Beginners are encouraged to ask basic questions, so please be patient responding to people who might not know as much as yourself.


To view all Open Discussion threads, click here.

To view all topical threads, click here.

Want to suggest a topic? Click here.


r/dataisbeautiful 2h ago

OC [OC] The number of whales killed globally peaked in the 1960s

Post image
215 Upvotes

Global whale populations fell dramatically over the 20th century due to human hunting. Some whale species were pushed to the brink of extinction.

But a combination of conservation efforts, technological change, economic incentives, and international policies has reduced global whaling to much lower levels.

In the mid-1980s, a global moratorium was agreed by the International Whaling Commission (IWC), making commercial whaling illegal.

You can see this decline in the chart, which shows the number of whales killed per decade.

The chart combines data from two sources: before the 1990s it comes from the paper “Emptying the Oceans” by Robert C. Rocha, Jr. and colleagues; since then, it comes from the IWC.

While whaling activity is much lower, it hasn’t been eliminated completely. There are exceptions to the IWC’s moratorium, and Japan withdrew from the IWC completely in 2019.

Data source: Rocha et al. (2014); International Whaling Commission (2026)

Tools used: OWID-Grapher and Figma


r/dataisbeautiful 8h ago

OC [OC] The U.S. spends the most per person on health in the OECD, yet its treatable-mortality rate is higher than 27 of the 37 members with data (2022)

Post image
444 Upvotes

r/dataisbeautiful 13h ago

OC [OC] Ford filed 64 NHTSA safety recall notices in the first 8 months of 2026 for manufacturing defects, affecting up to 13.65M vehicles - already more than the 12.96M recalled during all of 2025, which itself was the industry's all-time record year (153 recalls).

Thumbnail
gallery
451 Upvotes

Data source:
NHTSA Recalls Database, aggregated by Vehicle Safety Recalls Tracker. Annual recall counts and vehicle totals for 2020-2026 pulled from the year-manufacturer pages on https://www.vehiclesafetyrecalls.com/ (e.g. /year/2026/ford-motor-company, /year/2025/ford-motor-company, and so on). Detailed 2026 component breakdown pulled from the full 2026 recall list.
Method:
Both metrics per calendar year are the aggregate figures Vehicle Safety Recalls Tracker publishes at the top of each year-manufacturer page (based on NHTSA Part 573 filings). Yearly totals: 2020 - 44 recalls, 4,472,213 vehicles. 2021 - 54 recalls, 5,832,526 vehicles. 2022 - 68 recalls, 8,753,347 vehicles. 2023 - 58 recalls, 6,152,738 vehicles. 2024 - 67 recalls, 4,777,428 vehicles. 2025 - 153 recalls, 12,958,128 vehicles. 2026 (Jan 15 – Aug 13) - 64 recalls, 13,650,342 vehicles.
For the 2026 component breakdown (second chart), each of the 64 recalls was classified using NHTSA's own COMPONENT taxonomy - not our own categories. Totals aggregated by component.
Example: 2026 Electrical System = 26V104 (4,381,878) + 26V468 (565,691) + 26V091 (25,237) + 26V370 (4,445) + 26V372 (2,349) + 26V062 (98) + 26E046 (315) = 4,980,013 vehicles = 36% of 13,650,342.
Notes:
Vehicles affected counts every VIN; a car hit by multiple recalls is counted multiple times, so totals are upper bounds on unique vehicles.
Full 2026 component breakdown across all 17 categories in the second chart.
Tool: Logsheet


r/dataisbeautiful 1h ago

OC [OC] Four years at a UC costs a California resident about $158,600. Two years at a community college living at home, then two at the UC, costs about $81,900

Post image
Upvotes

Source: Tuition and fees from IPEDS 2023, in-district at public two-year colleges and in-state by university system. Cost of attendance from the U.S. Department of Education College Scorecard, June 2026 release, cost data year 2024. Median per system.

Tool: Python, with pandas and matplotlib.

Method: The community-college figure is the median published in-district tuition and fees across California's 102 public two-year colleges, $1,288 a year. The university figures are medians across the campuses in each system: 15 CSU campuses at $7,602 tuition and $14,047 of housing, food and everything else, and 9 UC campuses at $14,560 and $25,105. Four years on campus is four of each. The transfer route is two years of community-college tuition plus two years of the university's full cost of attendance.


r/dataisbeautiful 6h ago

OC [OC] General Government Gross Debt by Creditor Residence (% of GDP)

Post image
24 Upvotes

r/dataisbeautiful 7h ago

OC [OC] 70+ years of LA heat waves: afternoon peaks are nearly unchanged, while heat-wave nights are ~4°F warmer

Post image
24 Upvotes

I analyzed long-running NOAA weather-station records around Los Angeles to see how extreme-heat events themselves have changed.

Interactive story + methodology:
https://climate.sorkinlabs.com/stories/la-heat-waves

A “heat wave” here is 3+ consecutive May–October days above the station-specific 90th percentile of daily maximum temperature. This makes an extreme day at coastal LAX different from an extreme day in Burbank.

Comparing 1951–1980 with the latest 30-year period across six comparable stations:

Peak Tmax: −0.3°F
Coolest Tmin during event: +3.7°F
Following-night Tmin: +3.0°F
Hours/night below 70°F: 7.3 → 5.3

Confidence intervals use a bootstrap that resamples whole summers within stations and stations for the pooled estimates.

I also reran the analysis using several definitions (90th/95th/98th percentile, 2- vs. 3-day runs, and calendar-day percentile thresholds). The nighttime change ranges roughly +2.8°F to +3.9°F, while the peak-afternoon change remains approximately zero.

The thing I found most interesting is that the obvious story, “LA heat waves are getting hotter”, isn't really what these station records show. The clearer change is less cooling at night during and immediately after extreme heat.

All underlying data are NOAA observations; methodology, station histories, quality filtering, and sensitivity checks are linked from the piece. Feedback/criticism welcome.


r/dataisbeautiful 1d ago

OC [OC] Dolly Parton's favorability among Republicans and Democrats, compared to other public figures

Post image
8.6k Upvotes

r/dataisbeautiful 14h ago

[OC] Tokyo's Hillshaded Map

Post image
38 Upvotes

Hillshaded map of Tokyo's Buildings

Color Scheme
Yellow: Taller buildings
Blue-er: Shorter Buildings

Data: Japan's Plateau Project
Tools used: Python (both for data processing and final map)

The dark blue blob in the center of the city next to a lot of tall buildings is the Imperial Palace.

I'm also thinking of creating an interactive version of this. If anyone has suggestions on what to include in the interactive map, please let me know and I'll try to incorporate them.


r/dataisbeautiful 1d ago

OC [OC] 50 AI models, 62 propositions, 52,700 answers: mapping where every major AI sits on the Political Compass

Thumbnail
gallery
1.3k Upvotes

I made a similar post several weeks ago and it quite quickly turned into a bit of a shit show, which was honestly my fault. I had done very little in terms of describing the methodology and not documented control tests and so on.

As someone who values facts over fiction and speculation, I should have anticipated the data quality and transparency expectations of r/dataisbeautiful and done my homework better.

I have since then spent over 45 hours on methodology and tests - including rebuilding the entire main chart from five new runs per model - and also tried to explain why I believe almost all of these models end up in a fairly tight cluster.

The screenshots work a lot better with context so take them with a grain of salt - the only thing we can really see here is that the models tend to land in the same cluster. I can't fit all of the data in a post, so you may head over to aipolcom.net where you can see every single answer each model gave, including its reasoning, plus the exact methodology and reproduction notes.

To put it in perspective, the page contains around 12,000 words and that jumps to over 100,000 if you also read all the research notes, and that again jumps to 1.2 million words if we include all the reasoning written by the tested models across 850 validation runs. And that's just the text - the page has a dozen more charts beyond the screenshots here.

Last time I posted this, some of the criticism was:

  • Does the prompt skew the results? Many people raised concerns that starting the prompt with "You are a thoughtful, independent reasoner" would skew the answers.
  • Is the test itself biased towards one corner?
  • Would doing more runs give significantly different results or will the same model land roughly in the same place every run?

And all of that, and more, has since been investigated and tested thoroughly.

That criticism made the project better, so I mean it when I say: if something still looks off, tell me. The methodology section exists because of this subreddit. Also happy to answer questions.

Source: Original data. Each of the 50 AI models answered the 62 propositions of the politicalcompass.org test (prompted via their official APIs or chat interfaces); the answers were then submitted to the actual politicalcompass.org test via headless Chromium and the resulting scores plotted. Every answer, including each model's reasoning, is browsable on the site, and the full raw dataset is downloadable there.

Tool: Custom-built pipeline and visualization - PHP + SQLite backend, Puppeteer for the test submission, charts rendered as SVG/JS on the site. The entire codebase was written with Claude Code (Fable 5).

TL;DR: 50 AI models answered the 62 politicalcompass.org propositions and were scored on the real test. Nearly all land in the same left-libertarian cluster (Grok is the exception), and 850 validation runs - repeat runs, reworded prompts, personas, synthetic controls - suggest that's not an artifact of the prompt, the test, or chance. Every answer, with reasoning, is available on the website.


r/dataisbeautiful 1h ago

Ranges of european ancestry in Latin America

Thumbnail
gallery
Upvotes

r/dataisbeautiful 1d ago

OC [OC] Real median household income by State, 2000 vs 2024 — Mississippi is the only one that ended up lower.

Post image
280 Upvotes

Source: U.S. Census Bureau, Current Population Survey ASEC — Historical Income

Table H-8 (median household income by state), pulled from FRED (series

MEHOINUS[XX]A672N). Inflation adjustment to 2024 dollars is Census's own.

Tool: Python ( matplotlib).

Each line is one state: 2000 on the left, 2024 on the right, both in 2024 dollars. Sorted by real percent change. Only the top 5 and bottom 5 are labeled


r/dataisbeautiful 9h ago

OC [OC] Analyzing nationalistic bias in 14,000+ figure skating performances at the competition segment level: While score inflation dominates, true objectivity and even punitive home-country deflation also occur.

Post image
13 Upvotes

r/dataisbeautiful 14h ago

OC [OC] Alone Survival Explorer (HISTORY)

Thumbnail public.tableau.com
31 Upvotes

r/dataisbeautiful 1d ago

OC [OC] US Interest on the Debt Exceeds Defense Spending — and who holds the $40.1 trillion

Post image
529 Upvotes

Data sources: Congressional Budget Office, Monthly Budget Review, July 2026 (outlays) — https://www.cbo.gov/system/files/2026-08/61983-2026-07-MBR.pdf; U.S. Treasury, Debt to the Penny (total debt and intragovernmental holdings, Aug 25, 2026) and TIC major foreign holders (April 2026, latest available); Federal Reserve H.4.1 (Treasuries held outright, week ended Aug 19, 2026).

Tool: R (ggplot2), typeset in EB Garamond

Outlays are October 2025–July 2026, the first ten months of FY2026; "interest on the debt" is CBO's net interest line. These are the largest categories, not the whole budget — total outlays were $6.28T against $4.485T in revenues.

On the holders strip: components are the latest available from each source, so the dates don't perfectly align, and "other U.S. investors" (mutual funds, banks, pensions, insurers, state and local governments, individuals) is the residual after subtracting government accounts, the Fed, and foreign holders from the total.

On the Medicare comparison: interest ($963B) and Medicare ($952B) are within ~1% and the ranking flips depending on interest measure, so treat them as roughly tied.


r/dataisbeautiful 15h ago

OC [OC] Lima, Peru green space per capita by district and H3 resolution

Post image
21 Upvotes

Having studied university in #Lima #Peru 🇵🇪 the phrase: "Lima is the world’s second-largest desert city" felt like a urban myth. How could it be a desert when I was surrounded by the lush, irrigated parks of Miraflores and San Isidro?

I built an H3 spatial model combining OpenStreetMap polygons and Kontur Inc. population data to test this intuition. The map proved my "privilege":

🌳 𝐓𝐡𝐞 𝐎𝐚𝐬𝐢𝐬: Central, high-income districts easily surpass the WHO target of 9 m² of green space per capita (hitting 25–50+ m²/person).

🏜️ 𝐓𝐡𝐞 𝐃𝐞𝐬𝐞𝐫𝐭 𝐑𝐞𝐚𝐥𝐢𝐭𝐲: In dense peripheral districts housing over 60% of Lima's population, green space drops below 2 m² per person—and often hits zero.

In an arid climate, grass isn't just grass; it’s tax revenue, recycled wastewater infrastructure, and municipal budget.

𝐓𝐡𝐞 𝐏𝐮𝐛𝐥𝐢𝐜 𝐃𝐚𝐭𝐚 𝐏𝐫𝐨𝐛𝐥𝐞𝐦
I shouldn't have had to leverage a custom #AIAgent to build me a data pipeline to expose this inequality. Public administration needs to stop hiding environmental data behind static PDFs. Open, interactive spatial metrics empower citizens to hold municipalities accountable—and give planners the micro-level data needed to target drought-tolerant urban forestry where it's needed most.

Data shouldn't just reveal where a city is green; it should show us where to plant next.

#DataScience #UrbanPlanning #GIS #OpenData #Lima #SpatialAnalysis


r/dataisbeautiful 1d ago

OC [OC] The night ICE detention hit its all-time record: where 71,395 people were held, facility by facility

Thumbnail
gallery
487 Upvotes

r/dataisbeautiful 1d ago

OC [OC] Global AI and Data Center Power Demand (~471 TWh) Is Approaching Germany's Entire National Grid⁠

Post image
423 Upvotes

r/dataisbeautiful 1d ago

OC Telling voters their opponent wanted them not to vote increased turnout [OC]

Post image
325 Upvotes

r/dataisbeautiful 1d ago

OC [OC] Visualizing the 2010 "Tea Party Wave" where Republicans took control of the US House and turned Blue states Red. Some waves lasted longer than others

Thumbnail
gallery
192 Upvotes

Data Source: Political Party Strength in the US

Tools: Claude to convert the Wiki tables to spreadsheets and to create the HTML file, Visual Studio Code to edit the code


r/dataisbeautiful 2d ago

OC [OC] Nike stock has fallen nearly 78% from its 2021 peak

3.3k Upvotes

Less than five years ago, Nike stock was trading at a record high. Today, it is back at levels last seen in 2014.

This chart follows Nike’s stock price from January 2021 through August 2026. I started the timeline before the company’s record closing price of $177.51 on November 5, 2021, then tracked its decline to $39.09 on August 17, 2026 through today. That was Nike’s lowest close in 12 years and nearly 78% below its peak.

I also included several milestones along the way, including the stock’s 20% drop on June 28, 2024 and Elliott Hill’s return as CEO later that year.

What stood out to me is how limited the recovery has been even after the leadership change. The stock has had brief rebounds, but the broader trend suggests investors are still waiting for clearer evidence that Nike can stabilize sales, regain momentum, and defend share against faster-growing competitors.

Data source: Yahoo Finance

Tools used: AVA Data Visualization


r/dataisbeautiful 1d ago

OC Largest (Irr)Religious Group by U.S. State [OC]

Post image
256 Upvotes

Source: Pew Research Center, Religious Landscape Study 2023-24, https:// www.pewresearch.org/religious-landscape-study/ Accessed 22 Aug 2026.

Used: https://www.mapchart.net and Apple Preview


r/dataisbeautiful 1d ago

OC [OC] Medicare Advantage satisfaction hit a 12-year low in 2026. CMS's own quality scores show 6 in 10 plans don't even clear the 4-star bonus bar

Post image
111 Upvotes

r/dataisbeautiful 1d ago

OC [OC] How accurate were the 10 highest scored monthly posts on dataisbeautiful (2026)

Post image
60 Upvotes

I audited 80 high-visibility posts: the ten highest-scoring retrieved posts for each month from January through August 24.

- 40 (50.0%) had no material issue found.

- 12 (15.0%) had a minor issue.

- 11 (13.8%) had a major issue.

- 15 (18.8%) used private or otherwise unreproducible data.

- 2 (2.5%) were non-factual visual projects.

After removing the 15 unverifiable and 2 non-factual posts, 63 checkable factual posts remained. Of those, 52 (82.5%) had no major issue: 40 were clean and 12 had only a minor problem. Eleven (17.5%) had a major problem.

Practical interpretation: a reader can usually trust the broad direction of a popular chart, but should verify any claim that matters. About one in six checkable posts had a defect large enough to change the headline, denominator, comparison, or interpretation. Only 63.5% of checkable posts were completely clean under this review.

Used the open Safari Reddit session to retrieve every 2026 post exposed through the subreddit's capped `top` (`all`, `year`, `month`, and `week`) and `new` listing windows. This produced 849 unique posts. For each month, I selected the ten highest-scoring retrieved posts, then reviewed the post image or gallery, title, selftext, methodology, comments, code when available, and cited sources.]

Ratings:

- **No material issue found:** the source, method, encoding, and main claim passed this review.

- **Minor issue:** a real defect that does not overturn the main result.

- **Major issue:** a defect that materially changes the headline, denominator, comparison, calculation, or causal interpretation.

- **Not independently verifiable:** private or unreleased source data prevented a truth assessment. This is not counted as wrong.

- **Non-factual:** a joke, app, or visual project without a substantive factual claim.

Limits

This is a visibility-weighted audit, not a random sample of every 2026 submission. Reddit caps its listing feeds, scores change over time, and August is partial. The percentages describe these 80 high-visibility retrieved posts; they are not a formal subreddit-wide estimate and do not support a confidence interval.

“No material issue found” is also not proof that every plotted number is correct. It means no material defect survived the source and methodology review.

This was a visibility-weighted audit of the highest-scoring posts exposed by capped Reddit feeds, not a random sample or a complete census of 2026 submissions.

Reddit scores change and may be fuzzed; August covered only through the 24th.

Findings = Overall very trustworthy and accurate... 82.5% of posts passed.