Sign in

James MacGlashan

@jmac-ai.bsky.social
2.6K followers 1.5K following 1.4K posts

Ask me about Reinforcement Learning Research @ Sony AI AI should learn from its experiences, not copy your data. My website for answering RL questions: www.decisionsanddragons.com Views and posts are my own.

PostsRepliesMedia
James MacGlashan @jmac-ai.bsky.social · 3h
Honestly, good old classification losses without RL would probably be fine. And there are already a billion datasets for that. The exception is when you want it to be more of an agent where there are sequential dependencies. Then you probably need RL for that kind of competency.
030
James MacGlashan @jmac-ai.bsky.social · 3h
It's a salient question. My guess is it's not enough. The recipe for RL with LLMs is no secret and I'm betting what they're doing isn't that unique. RL on big LLMs is a bit of a moat in so far as it's hard for people to scale it up. Jev being more lean classification heads is probably easier.
110
James MacGlashan @jmac-ai.bsky.social · 3h
Ah fair enough, and thanks! While I have posted about that in a number of places, it's probably disperse. I'll think about writing a dedicated thread on it sometime!
000
James MacGlashan @jmac-ai.bsky.social · 5h
I agree it's a good idea. Zero-shot classification is pretty great for lots of things. But in addition to the big labs squashing it, I think it's also very amenable for small bespoke models which can eat its lunch too.
140
James MacGlashan @jmac-ai.bsky.social · 5h
It occurs to me you may have just been generally hoping to see an additional thread made on that rather than criticizing why I didn't cover it in this thread. If so, apologies! I have strived to cover that many times in the past, so I do think that's important!
100
James MacGlashan @jmac-ai.bsky.social · 5h
I regularly inform people that AI is a big field with many topics and methods. Even my profile description leans that way. I hope you’re not of the mind that every bluesky thread someone writes needs to include the totality of their positions. Because that would just be silly.
100
James MacGlashan @jmac-ai.bsky.social · 6h
As I stated in my original thread, there are a lot of valid criticisms that can be leveled at the companies making frontier models. I have made them myself many times. The claim that ML, or GenAI, isn't AI is not one of them. It's more productive to criticize the actual problems.
000
James MacGlashan @jmac-ai.bsky.social · 6h
Not only are those talked about, some of them are even used with LLMs. But even if you didn't see them talked about, that does not mean GenAI is not AI tech anymore than the Fast-forward graph planner is not AI tech because it doesn't use other areas of AI.
100
James MacGlashan @jmac-ai.bsky.social · 7h
To be clear, when I say nothing was stolen, I'm referring to your comment that the AI name was stolen. There are absolutely cases of pirated data that was stolen for training and the companies involved should be held accountable for those violations.
100
James MacGlashan @jmac-ai.bsky.social · 7h
GenAI and LLMs are very much still part of the AI field, developed by AI researchers. Nothing was "stolen." It is not correct that LLM development is unconcerned about correctness. Overhype does do damage, but GenAI remains AI technology that has considerable value.
100
James MacGlashan @jmac-ai.bsky.social · 08/10/2026
I mean, the RL is important, but it would be silly to to think we just won't extend the spaces its used. Less easily verified domains will make that harder to do, but by no means insurmountable.
0120
James MacGlashan @jmac-ai.bsky.social · 07/10/2026
Tweets are now battlefields.
140
James MacGlashan @jmac-ai.bsky.social · 07/10/2026
Now this is a risk that I worry about. And the US has the worst administration possible for avoiding it.
000
James MacGlashan @jmac-ai.bsky.social · 07/10/2026
I love how many people's first question from this (including my own) was: "Doom port when?"
0160
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
I appreciate your attempt to strongman it! :)
000
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
Of course they're fuzzy boundaries. Every social construct is. If fuzzy boundaries is too insurmountable for someone though, that should not lead to them to complaining that "ML is not AI" specifically. You need to embrace boundaries to make that claim or complain about terms altogether.
110
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
I believe modern spreadsheets now have methods developed by the ML field, in which case those specific parts are, yes. (Not sure, I avoid spreadsheets whenever possible.) Those features not developed by ML, and indirectly by AI, are not.
020
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
Yes, Neural Nets fell out of favor for a long time. This happens a lot. SVM were in vogue and then not. Same with Bayes Nets. or on the symbolic side, expert systems were the rage, then not. None of the ebbs and flows of what methods are in vogue means "Machine learning" isn't a subfield of AI.
110
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
Sorry, but a point in time where some subset of the population wanted to distance themselves because of funding difficulties in the 90s is not a compelling counterargument argument to the overwhelming history of the field and relationship that persists today. :p I appreciate the attempt though!
100
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
And like I said, if there is an anomaly or fib about its relationship it was anyone denying it in that time! That was the branding manipulation if anything was.
110
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
The "calculous" typo is going to haunt me forever.
030
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
Cross fertilization in academics is good and happens all the time in every field.
120
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
Those things absolutely were crucial to modern methods. But that doesn't change that it is an AI field. Discounting it as an AI field because it built on various math fields would be like denying statistics as its own field because it of its dependence on measure theory and calculous.
340
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
Our good ol' friend "the AI effect" :p
110
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
ML goes back to the late 50s. Even neural nets are an idea that started back then. Scientists may have preferred to focus on the ML name at some point to avoid funding issues, but it still remains inherently part of AI and its history. If there was a fib/anomaly, it was denying the relationship.
190
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
There is a mountain of things you can rightly criticize modern AI industry units of. This is not one of them and it is a bizarre thing to latch onto.
0260
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
Or you can go to the 1995 first edition of the Russell Norvig AI textbook, which includes a whole part dedicated to machine learning including decision trees, information theory, neural nets, RL, etc. www.mbit.edu.in/wp-content/u...
mbit.edu.in
2270
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
If you doubt it, you can read the wikipedia page where this is literally stated in multiple places in multiple ways. en.wikipedia.org/wiki/Machine...
Quote image of the wikipedia introduction on ML saying: "Machine learning (ML) is a field of study in artificial intelligence"Quote image of wikipedia history saying: "The term machine learning was coined in 1959 by Arthur Samuel, an IBM employee and pioneer in the field of computer gaming and artificial intelligence.[6][7] The synonym self-teaching computers was also used during this time period.[8][9]"Quote image of wikipedia ML section on relationship to AI. It includes an image of ML circle being contained in the AI cirlce. It also says: "As a scientific endeavour, machine learning grew out of the quest for artificial intelligence (AI). In the early days of AI as an academic discipline, some researchers were interested in having machines learn from data. They attempted to approach the problem with various symbolic methods, as well as what were then termed "neural networks"; these were mostly perceptrons and other models that were later found to be reinventions of the generalised linear models of statistics.[26] Probabilistic reasoning was also employed, especially in automated medical diagnosis.[27]: 488 "
1230
James MacGlashan @jmac-ai.bsky.social · 06/10/2026
Some people seem to think that ML wasn't originally considered AI and are criticizing labeling ML as AI as something recent. That's not right. ML is a subfield of AI. Not all AI is ML, but ML is AI. The history on this relationship is clear; they've always been associated this way.
57313
James MacGlashan @jmac-ai.bsky.social · 05/10/2026
This is an old article and not very good. The author doesn't seem to understand IP or how computers work and their proposed alternatives are in fact IP "alternatives." If you want to know why, this is a good essay on some of the things wrong with it. sergiograziosi.wordpress.com/2016/05/22/r...
sergiograziosi.wordpress.com
Robert Epstein’s empty essay
Sometimes reading a flawed argument triggers my rage, I really do get angry, a phenomenon that invariably surprises and amuses me. What follows is my attempt to use my anger in a constructive way, …
0204
James MacGlashan @jmac-ai.bsky.social · 05/10/2026
Yeah, this is an old article and is sadly pretty bad. It completely fails to make the case because the author does not apparently understand what IP is or how computers work. I found this critique, released back with the article, on point: sergiograziosi.wordpress.com/2016/05/22/r...
sergiograziosi.wordpress.com
Robert Epstein’s empty essay
Sometimes reading a flawed argument triggers my rage, I really do get angry, a phenomenon that invariably surprises and amuses me. What follows is my attempt to use my anger in a constructive way, …
020
James MacGlashan @jmac-ai.bsky.social · 05/10/2026
While I don't agree with the characterization of RL as simulating other data, if the real distinction @colin-fraser.net is after is the difference between a frozen probabilistic model (policy) and a comprehensive learning process, that is a meaningful distinction.
120
James MacGlashan @jmac-ai.bsky.social · 05/10/2026
An exception is if you meta-learn an RL algorithm inside the model. And to some extent LLMs do appear to capture aspects of that. But they're pretty limited as of now. There's a reason we still train new models instead of just unroll an old model with whatever meta-learning it captured.
120
James MacGlashan @jmac-ai.bsky.social · 05/10/2026
For sure there is more than RLHF, and I think that RLVR is much more important when it comes to capabilities. But it is true that it is a frozen model that is deployed to us. It's a policy, and a policy isn't the same thing as RL, it's an operand of it.
130
James MacGlashan @jmac-ai.bsky.social · 05/10/2026
Yep, I certainly agree that once you stop training, you lose the reward optimization! Naturally it doesn't have to be that way -- lots of people working on continual learning -- but right now it is and that's a limitation.
120
James MacGlashan @jmac-ai.bsky.social · 05/10/2026
Yes, that makes sense. If an oracle provides you optimal data, you should just do supervised training! But I still wouldn't describe RL as "a way to simulate different data," because without the oracle, you need to do more than simulate data. I think that is misleading, but we may otherwise agree.
110
James MacGlashan @jmac-ai.bsky.social · 04/10/2026
No one who has this take should ever complain about a manager in tech who doesn't understand it :p
050
James MacGlashan @jmac-ai.bsky.social · 04/10/2026
I think you mistyped 42.
000
James MacGlashan @jmac-ai.bsky.social · 03/10/2026
This is precisely the analogy I use for when it's useful!
040
James MacGlashan @jmac-ai.bsky.social · 03/10/2026
There's some gesturing to all-powerful rogue AI that I disagree with, but I think the employee's core concern is reasonable. Regardless of whether the worst doomer scenario is plausible, there remains lots of harm that can be caused by irresponsible actors even now. We need to fix that.
031
Reposted by James MacGlashan
NYU Center for Data Science @nyudatascience.bsky.social · 02/10/2026
Can LLMs introspect? Anthropic said yes. However, CDS PhD student Shashwat Singh, CDS Associate Professor Tal Linzen (@tallinzen.bsky.social) & CDS Faculty Fellow Shauli Ravfogel (@shauli.bsky.social) found the evidence falls short. nyudatascience.medium.com/cds-research...
nyudatascience.medium.com
CDS Researchers Challenge Anthropic’s Evidence That Language Models Can Introspect
In 2025, Anthropic reported that its Claude models could detect when researchers injected a concept directly into their neural activity…
13516
James MacGlashan @jmac-ai.bsky.social · 02/10/2026
Yeah much to my disappointment, the media has not done the community any favors with who they tend to cover!
110
James MacGlashan @jmac-ai.bsky.social · 02/10/2026
I do appreciate that :) But even setting me aside, I'm not sure I can think of anyone I collaborate with now or in the past 20 years for whom that description holds! Frontier labs and SV have caused a very distorted view of the AI community. But I do sadly get why that's the impression :/
110
James MacGlashan @jmac-ai.bsky.social · 02/10/2026
I guess I'm inclined to be defensive, but I think you at least need to add "in the bay area" as a qualifier :p It really doesn't hold well outside that in my experience.
110
James MacGlashan @jmac-ai.bsky.social · 02/10/2026
“Machines can be conscious in principle” does not imply that Claude is anymore than it implies that the chat bots of yore were, or any other software for that matter. If you want to make that claim, then you need to do a lot more work. And no, “it looks like it to me” is not doing the work.
070
James MacGlashan @jmac-ai.bsky.social · 01/10/2026
LLMs are fascinating because we kind of stumbled into these crucial ways to train them that have made them so powerful. Moreover, we still haven't applied these lessons everywhere in AI. I called it out, because it's important that we apply the lessons to more! :)
1120
James MacGlashan @jmac-ai.bsky.social · 01/10/2026
I brought this up because a lot of people have gotten into practical trouble because they stop at "my model can in principle represent this optimal solution." That isn't enough because even if your model can, that doesn't mean it can be found by your training process.
1181
James MacGlashan @jmac-ai.bsky.social · 01/10/2026
RL is often useful because it provides a means to organically discover the algorithmic steps even if it wasn't explicit in the data. But as I described above, you can also do it in more supervised ways. LLM training tends to have both (RL and algorithm traces in datasets).
1131
James MacGlashan @jmac-ai.bsky.social · 01/10/2026
This is what makes LLMs much more than fixed-capacity function approximation -- they can do general compute if you change the format to allow for intermediate steps _and_ have a process to train those steps.
1143
James MacGlashan @jmac-ai.bsky.social · 01/10/2026
With that change, now you're in business. The encoding for each step in the weights, given the observation of them, is much more tractable for SGD to find. Even better, it can now handle _indefinitely large inputs_ because it can unroll the algorithm it learned to arbitrary length.
191