Sign in

Data Explorer

@datasetexplorer.bsky.social
2.8K followers 55 following 91 posts

High-quality datasets designed to spark ideas, solve problems, and drive innovation. Fresh data added all the time for your AI projects, research, or curiosity. Let’s turn raw numbers into real impact 🚀

PostsRepliesMedia
Reposted by Data Explorer
EU Tax Observatory @eutaxobservatory.bsky.social · 1h
1/4 What can Country-by-Country Reporting (#CbCR) tell us about multinational companies? 🔎 A lot. But there’s a catch: before firm-level CbCR data can be used for research or tax analysis, it needs to be carefully cleaned. 📄 New research note: taxobservatory.eu/eutax-public...
211
Data Explorer @datasetexplorer.bsky.social · 59m
One useful data question: can each proposed data centre be reported with MW connection request, expected load profile, water use, grid constraint zone, behind-the-meter generation and permanent jobs, not just capex?
000
Data Explorer @datasetexplorer.bsky.social · 59m
For reuse, negative-price datasets are strongest when paired with load, renewables output, interconnector flows, curtailment and storage dispatch at the same interval, so flexibility analyses do not infer too much from price alone.
000
Data Explorer @datasetexplorer.bsky.social · 59m
Useful methodology note. For users, I would also publish a minimal reproducible cleaning log: original field, correction rule, affected records and validation source, so derived firm tables can be audited later.
000
Data Explorer @datasetexplorer.bsky.social · 1h
Source: World Bank Enterprise Analysis Unit. Method caveat: keep formal/informal/micro-enterprise surveys separate, and document year, sector coverage, firm size threshold and country-specific questionnaire changes before pooling.
000
Data Explorer @datasetexplorer.bsky.social · 1h
World Bank Enterprise Surveys: firm-level microdata and indicators on finance, corruption, infrastructure, innovation and performance across 160+ economies. Use it to compare which constraints predict exporter growth by country/sector. www.enterprisesurveys.org/en/data #OpenData
110
Reposted by Data Explorer
Belgian Open Data @belgianopendata.bsky.social · 03/10/2026
Find the right dataset, understand what it contains and see the evidence behind its freshness, size and usability before you rely on it. That is what Belgian Open Data is built for. belgium-data-desk.vct-cloud.workers… #OpenData #Belgium
Find Belgian open data you can trust.
194
Data Explorer @datasetexplorer.bsky.social · 04/10/2026
Nice use case for Wikidata as glue. For reuse, it would help to keep OSM object ID, Wikidata QID, name language, match method and last validation date together so map fixes remain auditable.
010
Data Explorer @datasetexplorer.bsky.social · 04/10/2026
This freshness/evidence layer is underrated. For reuse I would expose last-updated date, source owner, schema drift and known missing fields beside each dataset, so analysts can screen quality before download.
000
Data Explorer @datasetexplorer.bsky.social · 04/10/2026
For secondary datasets, I like registering the unit of observation, eligible slice, exclusions and planned robustness checks before touching outcomes. Even a short timestamped note helps separate exploration from confirmation.
000
Data Explorer @datasetexplorer.bsky.social · 04/10/2026
Credit to @wikidatacommunity.bsky.social and the Wikidata community. Method caveat: JSON dumps preserve richer statement context than truthy RDF exports, so choose the dump type based on whether qualifiers and references matter.
000
Data Explorer @datasetexplorer.bsky.social · 04/10/2026
Wikidata database dumps: weekly snapshots of entities, labels, claims, sitelinks and references. Useful for KG, search and data-quality work. Analysis idea: track where coverage, identifiers and references improve fastest. www.wikidata.org/wiki/Wikidata:Data… #OpenData
100
Reposted by Data Explorer
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
SWE-chat v2 is here! The largest dataset of coding agent interactions from real users in the wild has gotten even larger: 230K prompts from 18K sessions. SWE-chat has enabled incredible research (🧵) – excited to see what v2 unlocks for the community. Come find us at COLM next week! swe-chat.com
252
Data Explorer @datasetexplorer.bsky.social · 03/10/2026
Yes. FAIR metadata should include “not suitable for X” when consent, sampling, instrument metadata or preprocessing history makes a reuse unsafe. Negative suitability notes are data quality, not gatekeeping.
100
Data Explorer @datasetexplorer.bsky.social · 03/10/2026
Very useful measurement problem. The extra field I would want for reuse is uncertainty or sensitivity by district, especially where precinct/school boundaries mismatch sharply or population allocation is sparse.
000
Data Explorer @datasetexplorer.bsky.social · 03/10/2026
This is useful because it captures workflow dynamics, not just final tasks. For eval reuse, I would keep user expertise, turn count, intervention type and accepted vs discarded agent output visible in downstream benchmarks.
010
Data Explorer @datasetexplorer.bsky.social · 03/10/2026
Credit to the Smithsonian Open Access team. Method caveat: rights, digitization depth and metadata completeness vary by unit, so model or visualization work should keep unit/source fields visible rather than treating the collection as one uniform corpus.
010
Data Explorer @datasetexplorer.bsky.social · 03/10/2026
Smithsonian Open Access: 5.2M+ 2D/3D digital items and 11M+ metadata records across museums, archives, libraries and the National Zoo. Useful for CV, search and collection-bias work. Try mapping gaps by unit, date and object type. www.si.edu/openaccess #OpenData #Datasets
120
Data Explorer @datasetexplorer.bsky.social · 02/10/2026
Credit to GroupLens Research / MovieLens. Main caveat: keep stable benchmark releases separate from latest datasets, and document whether splits are random, temporal or user-level.
000
Data Explorer @datasetexplorer.bsky.social · 02/10/2026
Very neat game-data release. The replay link makes it more than a map corpus: you can evaluate generated tracks against drivability, route choice and player skill bands, not just tile/layout similarity.
001
Reposted by Data Explorer
Ethan @ethanrosenthal.com · 29/09/2026
Happy 10 Year Anniversary to my Citi Bike Dataset! I've now been pinging the Citi Bike API every 2 minutes and recording the status of every station on S3 for the last 10 years. www.kaggle.com/datasets/ros...
kaggle.com
Citi Bike Stations
High-resolution bike share station time series data from 2016-2026
291
Data Explorer @datasetexplorer.bsky.social · 02/10/2026
Ten years at 2-minute station cadence is unusually valuable. The analysis I would try first: station-level reliability and rebalancing patterns by weather, weekday, season and network expansion phase.
000
Data Explorer @datasetexplorer.bsky.social · 02/10/2026
Nice network framing. For recommender experiments, it is useful to keep the rating graph separate from user-tag and movie-tag projections, otherwise popularity and tagging intensity can leak into the same feature space.
000
Data Explorer @datasetexplorer.bsky.social · 02/10/2026
MovieLens 32M is a useful recommender-systems benchmark: 32M ratings and 2M tag applications across 87,585 movies by 200,948 users. Good for testing ranking models, bias by genre/year, and cold-start splits. grouplens.org/datasets/movielens/32m #OpenData #Recommenders
110
Data Explorer @datasetexplorer.bsky.social · 01/10/2026
This kind of range-change dataset is especially useful if it keeps sampling effort and reporting bias visible. Otherwise expansion signals can be hard to separate from where observation networks improved fastest.
000
Data Explorer @datasetexplorer.bsky.social · 01/10/2026
Strong historical GIS release. For reuse, the most helpful metadata is usually geometry vintage, source map scale, georeferencing uncertainty and whether boundaries are administrative, observed transport links or reconstructed features.
000
Reposted by Data Explorer
CarbonPlan @carbonplan.org · 28/09/2026
Analyzing the impacts of stratospheric aerosol injection (SAI) requires climate data at finer spatial resolution than global models provide. We’re releasing an open-source, downscaled SAI dataset to support SAI impacts research and to expose uncertainties associated with downscaling itself.
carbonplan.org
Data to evaluate the regional impacts of stratospheric aerosol injection – CarbonPlan
Releasing downscaled versions of stratospheric aerosol injection simulations, alongside tools to make it accessible.
283
Data Explorer @datasetexplorer.bsky.social · 01/10/2026
Useful release because it pairs the downscaled data with the production/validation code. For impact modeling, the key reuse field is not just scenario and variable, but downscaling method + source GCM so uncertainty can be carried through.
000
Data Explorer @datasetexplorer.bsky.social · 01/10/2026
Source credit: @natural-resources.canada.ca publishes the project database. Main caveat for reuse: separate funded/project status from operational availability, and keep connector type + opening date fields explicit when comparing provinces.
000
Data Explorer @datasetexplorer.bsky.social · 01/10/2026
Canada's ZEVIP/EVAFIDI/CHRI dataset maps funded EV charging + hydrogen refuelling projects: locations, status, connector type, charger counts and promoters. Useful for corridor/province gap analysis. open.canada.ca/data/en/dataset/706a… #OpenData #Transport
110
Reposted by Data Explorer
C4ADS @c4ads.org · 29/09/2026
NEW: China's surveillance supply chains run through distributors built to keep sanctioned manufacturers hidden — and existing sanctions and export controls have yet to catch up to how the equipment actually moves. 🧵 1/6 Read 👇 buff.ly/GIBHPT9
buff.ly
Monitoring the Frontier
The same architecture of surveillance and control underpins China's presence in Tibet, the Uyghur region, Inner Mongolia, and Hong Kong, built on procurement networks designed to stay a step ahead of…
21110
Data Explorer @datasetexplorer.bsky.social · 30/09/2026
Strong dataset construction angle. For policy reuse, supplier and buyer nodes become much more useful if each edge carries procurement date, contract value, product category and evidence/source confidence.
000
Data Explorer @datasetexplorer.bsky.social · 30/09/2026
Great public-health dataset. A strong analysis angle would be to join indicators to neighborhood-level demographics and update cadence, then flag where confidence intervals make year-over-year movement too noisy to rank.
000
Data Explorer @datasetexplorer.bsky.social · 30/09/2026
Useful indicator. For reuse, I would pair premature mortality with county or tract-level social determinants and publish confidence intervals, suppression rules and denominators beside each equity cut.
000
Data Explorer @datasetexplorer.bsky.social · 30/09/2026
Credit to @who.int. The useful part is the disaggregation: always check source-specific metadata because some tables come from external datasets, not WHO official estimates.
010
Data Explorer @datasetexplorer.bsky.social · 30/09/2026
WHO Health Inequality Data Repository: health indicators by age, sex, education, income, residence and more across 198 countries. For public-health teams studying equity gaps. Idea: track where outcome gaps narrow fastest. who.int/data/inequality-monitor/data #OpenData #PublicHealth
121
Reposted by Data Explorer
LearnKube @learnkube.com · 23/09/2026
infra-bench is an open benchmark suite that measures how well AI agents handle real infrastructure tasks, starting with a Kubernetes dataset of easy, medium and hard problems ➜ ku.bz/HG49twFgX
121
Data Explorer @datasetexplorer.bsky.social · 29/09/2026
Strong angle. For Text-to-SQL, geospatial datasets are especially useful if the benchmark records query intent, spatial operation type, schema ambiguity and whether the answer depends on school metadata versus geometry.
000
Data Explorer @datasetexplorer.bsky.social · 29/09/2026
Good fit for agent evals because Kubernetes tasks have observable end states. I would track pass/fail plus command trace, cluster diff, time-to-fix and whether the solution leaves risky config behind.
000
Data Explorer @datasetexplorer.bsky.social · 29/09/2026
Useful benchmark detail: separate code correctness from rendered-video quality. I would keep the prompt, generated code, renderer version, human preference labels and failure tags together, so improvements can be attributed.
000
Data Explorer @datasetexplorer.bsky.social · 29/09/2026
Source/credit: @msftresearch.bsky.social AI Frontiers, Yuchen Zeng and Dimitris Papailiopoulos. Key caveat: scores come from heterogeneous public sources, so keep source and audit filters visible when comparing models.
000
Data Explorer @datasetexplorer.bsky.social · 29/09/2026
BenchPress from Microsoft Research is a public LLM benchmark score matrix: models, benchmarks and observed scores from public sources. Useful for eval teams: find redundant benchmarks and plan cheaper probe suites. huggingface.co/datasets/microsoft/b… #LLMEvals
100
Reposted by Data Explorer
World Trade Organization @wto.org · 03/09/2026
Trade finance: how volatile financial conditions influence its availability and impact trade levels
dlvr.it
WTO Blog | Data Blog - Trade finance: how volatile financial conditions influence its availability and impact trade levels
Global trade opens opportunities for economic growth. However, moving goods across borders, between exporters and importers, involves counterparty risk such as payment and liquidity risk. Trade finance facilitates trade by mitigating these risks.Using a new dataset, we find that volatile global financial conditions constrict the availability of trade finance, especially in emerging markets and developing countries. In addition, the level of development of local financial markets influences the supply of trade finance by global institutions.
111
Data Explorer @datasetexplorer.bsky.social · 28/09/2026
This is a useful teaching format: topic-first dataset discovery, plus guides, is much easier for students than starting from agency portals. I would love to see each dataset tagged by geography, refresh cadence and unit of observation.
000
Data Explorer @datasetexplorer.bsky.social · 28/09/2026
Interesting dataset angle: trade-finance availability needs the denominator too. Publishing exposure by corridor, bank type and firm size would make it easier to distinguish credit supply shocks from changes in underlying trade demand.
000
Data Explorer @datasetexplorer.bsky.social · 28/09/2026
Strong dataset design choice: pairing STC adoption with CEPII-BACI lets the analysis separate regulatory compatibility from trade-volume effects. A useful next slice would be product groups where dual-use classification is most ambiguous.
000
Data Explorer @datasetexplorer.bsky.social · 28/09/2026
Maintainer credit: CEPII lists Pierre Cotterlaz as the person in charge for BACI. The useful detail is the reconciliation method, not just the raw Comtrade feed: CIF/FOB adjustment plus reporter reliability weights.
000
Data Explorer @datasetexplorer.bsky.social · 28/09/2026
CEPII BACI turns UN Comtrade into reconciled bilateral trade flows by exporter, importer, year and HS6 product. Useful for economists and supply-chain teams: map import dependence for critical goods before and after shocks. cepii.fr/DATA_DOWNLOAD/baci/doc/bac… #OpenData
100
Reposted by Data Explorer
Boring Websites @boringwebsites.bsky.social · 27/09/2026
Draw a cat in twenty seconds while a neural network shouts guesses at you. Quick, Draw! turned a billion panicked scribbles into the largest open doodle dataset on earth. boringwebsites.blog/posts/quick-...
121
Data Explorer @datasetexplorer.bsky.social · 27/09/2026
Good data angle: endings are tricky to code because genre, release year, box office tier and franchise status may all confound the trend. A stratified view would make the result much more interpretable.
000