Sign in

Chao Wang

@excel-wang.bsky.social
25 followers 20 following 87 posts

Associate Professor in health and social care statistics at Kingston University. PhD in econometrics.

PostsRepliesMedia
Chao Wang @excel-wang.bsky.social · 01/10/2026
🚨BREAKING. President Trump has renamed "United States" to "Unites States".
000
Chao Wang @excel-wang.bsky.social · 30/09/2026
Poor name choice. Have they never Googled their proposed model name?
000
Chao Wang @excel-wang.bsky.social · 29/09/2026
Given the nature of LLMs, they also have limitations in other scenarios, such as driving. The human brain can write text, drive a car, and play games like chess or Go, demonstrating a more versatile kind of intelligence.
001
Chao Wang @excel-wang.bsky.social · 29/09/2026
LLMs are essentially next-token prediction. It tells us, given past game records, which move is most likely to happen. AlphaGo on the other hand, tells us, given the possible moves, which move has the best chance of winning.
100
Chao Wang @excel-wang.bsky.social · 29/09/2026
In fact, when it comes to Go, humans have actually managed to catch up. www.kedglobal.com/artificial-i...
kedglobal.com
Go grandmaster Shin defeats AI KataGo in historic human victory - KED Global
Shin Jin-seo, the world's top-ranked Go player, on Tuesday completed a dramatic comeback against the world’s premier artificial intelligence Go engine, K
100
Chao Wang @excel-wang.bsky.social · 29/09/2026
If you think LLM is the way to achieve general intelligence (AGI?), think again. Using chess/Go as an example, LLMs are still far behind top human players, let alone top chess/Go AIs such as AlphaGo (which is not an LLM). x.com/i/grok/share...
x.com
LLMs Poor at Chess and Go
General-purpose LLMs are mediocre at chess (roughly beginner to intermediate club level at best) and very weak at Go. They cannot beat professional human players or top specialized AIs like Stockfish,...
100
Chao Wang @excel-wang.bsky.social · 23/09/2026
This is best explained by the researcher in the YouTube video (from 15:00). I initially thought Study 2 was weak as well.
110
Chao Wang @excel-wang.bsky.social · 23/09/2026
Study design does not necessarily dictate the level of evidence, as often depicted in "evidence pyramid". Context is important. See the example below (youtu.be/3pxudE0GNAo?...). Many people thought Study 2, a cross-sectional study, is the weakest. However Study 2 is arguably the strongest.
110
Chao Wang @excel-wang.bsky.social · 14/09/2026
Interesting article. It shows LLM cannot discover causal relationships purely based on the data without experts’ input.
0112
Chao Wang @excel-wang.bsky.social · 10/08/2026
I doubt legal fees and fines will be anywhere significant compared to NHS' overall budget. You don't think it's a great use of her resources because you don't agree with her.
100
Chao Wang @excel-wang.bsky.social · 06/08/2026
This reminds me the story about whether Jaffa Cake is a cake or biscuit www.crunch.co.uk/knowledge/ar...
crunch.co.uk
Jaffa Cakes - Cake or Biscuit? VAT's the Difference | Crunch
Is a Jaffa Cake a cake, or a biscuit? It took a case in court to come up with a full answer. Find out the answer here.
000
Reposted by Chao Wang
Jess Calarco @jessicacalarco.com · 31/07/2026
For example: When I teach qualitative methods, I ask my students which of these foods are sandwiches, and there's always lots of disagreement. So, we talk about how they can explain their criteria and try to persuade others to agree with them. But they can't objectively prove they're right.
Six images of food items: tacos, wrap, hot dog, banh mi, avocado toast, burger.
11566
Reposted by Chao Wang
Adam Kucharski @adamjkucharski.bsky.social · 05/08/2026
Two fictional (presumably AI) stories in recent X posts, now repeated as a Google summary. No need to tell tall tales about your achievements when the AI ecosystem will now do it for you…
65119
Chao Wang @excel-wang.bsky.social · 03/08/2026
In reality it’s more like “we can't possibly do that”… until we could 🤔
100
Chao Wang @excel-wang.bsky.social · 31/07/2026
I’m not sure about that. I think @jmwooldridge.bsky.social demonstrated Poisson regression doesn’t have any distribution assumptions.
281
Chao Wang @excel-wang.bsky.social · 14/07/2026
Also this headline "fatal/serious injuries ⬇️ 34%" is hugely misleading. The 34% headline figure refers to the intervention group alone, not being compared with the control group! It's disingenuous to present 34% as the effect of 20mph!
000
Chao Wang @excel-wang.bsky.social · 14/07/2026
Where is the evidence that they are "negligible" (I'd be keen to read the full economic analysis), or it even "saves lives"? This non-peer-reviewed "study" by TfL which pushes this policy (so conflict of interest)? The latest meta-analysis shows otherwise. bsky.app/profile/exce...
110
Chao Wang @excel-wang.bsky.social · 07/07/2026
You will get a better chance with closed-source feature-rich (i.e. not heavily relying on third-party packages to do anything useful) software.
010
Chao Wang @excel-wang.bsky.social · 02/07/2026
main problem people had with this blog was the author seemed to suggest collecting more data/variables is the main solution to address confounding in observational study.
210
Chao Wang @excel-wang.bsky.social · 02/07/2026
I agree that everything else being equal (e.g. same study design; same modelling approach), having more variables is good. But rarely are things equal in reality. Collecting more data can be challenging, and designs like Regression Discontinuity and DiD are much less data hungry. I think the...
110
Chao Wang @excel-wang.bsky.social · 02/07/2026
Also the idea that you cannot control for unobserved confounders is a myth. You can do it using methods like instrumental-variable regression (e.g. Mendelian randomisation study), latent class models, difference-in-difference models in the case of panel data...
100
Chao Wang @excel-wang.bsky.social · 02/07/2026
Yes you can leave out colliders (and instruments) but this "just collecting more covariates so RCT is no longer needed" approach actually makes the analysis harder not easier, as you have to go through the long list of covariates to decide which to include or exclude.
100
Chao Wang @excel-wang.bsky.social · 02/07/2026
So yes the more covariates we have, the higher chances we catch all confounders (this part is true); but it also means the higher chances we catch collider and instrument variables... then you need to think carefully which covariates to adjust.
000
Chao Wang @excel-wang.bsky.social · 02/07/2026
Adjusting collider and instrument variables can lead to bias link.springer.com/article/10.1... so it is not necessarily true that the more covariates in a model, the better.
link.springer.com
Principles of confounder selection - European Journal of Epidemiology
Selecting an appropriate set of confounders for which to control is critical for reliable causal inference. Recent theoretical and methodological developments have helped clarify a number of principle...
210
Chao Wang @excel-wang.bsky.social · 11/06/2026
"The study was observational"... Sure but how can an RCT prove "direct link"? Obviously unethical but imagine somehow you did an RCT on smacking, how can it prove smacking "directly" affects GCSE grades but not through some mediator?
120
Chao Wang @excel-wang.bsky.social · 09/06/2026
The abstract says the study covers 2007-2011? The App Store was launched in 2008. Apple itself boasted about its cultural impact: "it ignited a cultural, social and economic phenomenon that changed how people work, play, meet, travel and so much more." www.apple.com/uk/newsroom/...
apple.com
The App Store turns 10
When Apple introduced the App Store on July 10, 2008 with 500 apps, it ignited a cultural, social and economic phenomenon.
010
Chao Wang @excel-wang.bsky.social · 26/05/2026
Worth noting you were likely testing the consumer version of Copilot, not M365 Copilot that Adam was using (I know it's confusing). There is no "smart" mode but "auto" mode in M365 Copilot. The GPT 5.5 Think Deeper mode works fine. bsky.app/profile/exce...
010
Reposted by Chao Wang
Chao Wang @excel-wang.bsky.social · 21/05/2026
Tried a few models and both Gemini 3.5 Flash (Extended, screenshot1) and M365 Copilot GPT 5.5 (Think Deeper, s2) got the correct answer. Some reported Claude Opus 4.7 is also fine. Yes more advanced mode may not always be better, but they tend to be more correct than instant/quick/... mode.
001
Chao Wang @excel-wang.bsky.social · 21/05/2026
Tried a few models and both Gemini 3.5 Flash (Extended, screenshot1) and M365 Copilot GPT 5.5 (Think Deeper, s2) got the correct answer. Some reported Claude Opus 4.7 is also fine. Yes more advanced mode may not always be better, but they tend to be more correct than instant/quick/... mode.
001
Chao Wang @excel-wang.bsky.social · 18/05/2026
While John's data visualisation work is really impressive, it's really far from ideal to establish causality.
000
Chao Wang @excel-wang.bsky.social · 12/05/2026
The key question is whether the subsample (elite players in this case) can be a distinct "class" (as in a latent class analysis). If yes, then Simpson's paradox (en.wikipedia.org/wiki/Simpson... ) comes into play.
en.wikipedia.org
Simpson's paradox - Wikipedia
000
Chao Wang @excel-wang.bsky.social · 11/05/2026
Nice benchmark "Which LLM writes the best Stata code?" www.khaledeltokhy.com/benchmarks/
000
Chao Wang @excel-wang.bsky.social · 01/04/2026
"These Terms don’t apply to Microsoft 365 Copilot apps or services unless that specific app or service says that these Terms apply."
000
Chao Wang @excel-wang.bsky.social · 01/04/2026
"These Terms don’t apply to Microsoft 365 Copilot apps or services unless that specific app or service says that these Terms apply." You do realise they have a separate Microsoft 365 Copilot app right?
000
Reposted by Chao Wang
Sarah O'Connor @sarahoconnorft.ft.com · 26/03/2026
In spite of all the talk of Claude Code and Codex meaning the end of humans writing code, software job adverts are actually going up, according to @jburnmurdoch.ft.com's crunching of millions of job ads for this week's The AI Shift www.ft.com/content/7325...
20564174
Chao Wang @excel-wang.bsky.social · 25/03/2026
Based on my understanding of the paper above: AUC, O:E ratio, and calibration slope (and possibly some R2 measures).
110
Chao Wang @excel-wang.bsky.social · 25/03/2026
Probably due to the separation of powers. In the parliamentary system, the party (or parties) which controls the parliament forms the government ("administration"). So they are both lawmakers and "administrators".
010
Chao Wang @excel-wang.bsky.social · 25/03/2026
There are no assessments on calibration? Also F1 is not recommended. www.thelancet.com/journals/lan...
thelancet.com
Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
Numerous measures have been proposed to illustrate the performance of predictive artificial intelligence (AI) models. Selecting appropriate performance measures is essential for predictive AI models i...
100
Chao Wang @excel-wang.bsky.social · 01/03/2026
bsky.app/profile/exce...
010
Chao Wang @excel-wang.bsky.social · 26/02/2026
Use LLMs to automate statistical reports drwangstatsconsulting.wordpress.com/2026/02/26/u...
drwangstatsconsulting.wordpress.com
Use LLMs to automate statistical reports
Large Language Models (LLMs) can enhance productivity by automating code generation for statistical reports. Using Gemini and Stata, users can input analysis code, which Gemini converts to Stata co…
000
Reposted by Chao Wang
Kevin Mitchell @wiringthebrain.bsky.social · 21/02/2026
The strongest version of this illusion I’ve seen! Absolute head-wrecker!
25388122
Chao Wang @excel-wang.bsky.social · 30/01/2026
Inflation <-> wage growth. No? en.wikipedia.org/wiki/Wage%E2...
en.wikipedia.org
Wage–price spiral - Wikipedia
000
Chao Wang @excel-wang.bsky.social · 30/12/2025
🚨Breaking: free stuff is adopted faster than something you have to pay for.
010
Chao Wang @excel-wang.bsky.social · 09/12/2025
The HRT-on-heart disease story tells us it is quite often that discrepancies between observational studies and RCTs are not due to difference in designs but they are in different contexts and thus essentially answering different questions. www.sciencedirect.com/science/arti...
sciencedirect.com
The HRT controversy: observational studies and RCTs fall in line
010
Reposted by Chao Wang
Jason Concepcion @netw3rk.bsky.social · 06/12/2025
Couldn’t Sam Altman just ask ChatGPT how to make itself profitable
684610762
Reposted by Chao Wang
Thomas House @tah-sci.com · 24/11/2025
This by @whippletom.bsky.social is brilliant. A little dose of epistemic humility goes a long way. www.thetimes.com/comment/colu...
thetimes.com
We’ll need good data next time or lockdown arguments multiply
At the start of the pandemic two professors, Martin Landray and Peter Horby, did something that should have been banal but was also rare: they tested drugs
074
Reposted by Chao Wang
Darren Dahly @statsepi.bsky.social · 24/10/2025
We tried to tell y'all to stop calling everything "AI" many years ago and you just wouldn't listen and now the poor machine learners must also suffer alongside the statisticians 😜
68911
Chao Wang @excel-wang.bsky.social · 18/11/2025
P-value has been severely criticised for 7 decades (the paper in the screenshot was published in 1994). I guess it's not going to go away anytime soon...
010
Chao Wang @excel-wang.bsky.social · 17/11/2025
One can certainly view it as an intercept-only linear regression model. One can also view it as a nonparametric moving-average smoother with the purpose of smoothing noise and giving a baseline, which does not attempt to model the data-generating process.
010
Chao Wang @excel-wang.bsky.social · 17/11/2025
Yes but this could either mean "the deaths are expected to continue to decrease" (without pandemic) or "the deaths are expected to increase due to 'dry tinder effect'" (www.sciencedirect.com/science/arti...). A simple average is not perfect but it's more model-free.
100