Sign in

Gregory Marton

@gregory-marton.bsky.social
554 followers 1.3K following 64 posts

GenAI adjunct at Tufts, ft dad, cs tutor. www.seidellmarton.us/gremio www.linkedin.com/in/gregory-marton

PostsRepliesMedia
Reposted by Gregory Marton
Xiaoyan Bai @elenal3ai.bsky.social · 27/05/2025
🚨 New paper alert 🚨 Ever asked an LLM-as-Marilyn Monroe who the US president was in 2000? 🤔 Should the LLM answer at all? We call these clashes Concept Incongruence. Read on! ⬇️ 1/n 🧵
13017
Reposted by Gregory Marton
Sung Kim @sungkim.bsky.social · 26/05/2025
Do LLMs Think Like Humans? They find that, While LLMs achieve broad categorical alignment with human judgment, they falter in capturing fine-grained semantic nuances such as typicality and, critically, exhibit vastly different representational efficiency profiles.
44510
Reposted by Gregory Marton
Neil Traft @ntraft.bsky.social · 01/04/2025
Kudos to @nytimes.com for covering ARC-AGI in such an exquisite example of interactive data journalism. Amazing spot for @fchollet.bsky.social as well. www.nytimes.com/interactive/...
nytimes.com
Are You Smarter Than A.I.?
Some experts predict that A.I. will surpass human intelligence within the next few years. Play this puzzle to see how far the machines have to go.
1406
Reposted by Gregory Marton
Ethan Mollick @emollick.bsky.social · 26/03/2025
The funny thing about multimodal image generation as released in the last week by Google and OpenAI is that now LLM image generation works like how most people using LLMs for the past two years always thought LLM image generation works.
1776
Reposted by Gregory Marton
Ethan Mollick @emollick.bsky.social · 13/03/2025
If you have used LLM image generators, you know they are hard to control: LLM had to send a prompt to a separate image generation tool, it did not not make the image. Gemini is the first public release of a full multimodal LLM that can directly make images. This allows the systems to do detail work
512119
Reposted by Gregory Marton
Michael Friendly @datavisfriendly.bsky.social · 09/03/2025
📊 #dataviz Putting the people 👨‍👨‍👧‍👦back into charts that talk about them. An interesting histogram of numeracy scores for U.S. vs. some other countries, using Alberto Cairo's [weepeople font](github.com/propublica/w...) to show the people involved in these distributions. Src: bit.ly/3FrDq0v 👇
A histogram showing the distribution of numeracy scores for students in the U.S. vs. other benchmark countries. The histogram uses an icon array made up of little figures of people, using the `weepeople` font.
The chart has the title "U.S. Numeracy Education has room for improvement"
3237
Reposted by Gregory Marton
Melanie Mitchell @melaniemitchell.bsky.social · 09/03/2025
As part of the Science Homecoming project I wrote an opinion piece for the Albuquerque Journal on the dangers of cuts to federal science funding: www.abqjournal.com/opinion/arti... @sciencehomecoming.bsky.social @cantlonlab.bsky.social @spiantado.bsky.social
abqjournal.com
COLUMN: Science funding cuts threaten economy
If you’ve ever been treated for a medical problem, used a cellphone or a computer, or been excited by robots exploring Mars or gene-editing therapy, then your life has been
012743
Reposted by Gregory Marton
Tom Pepinsky @tompepinsky.com · 08/03/2025
The public really needs to understand this. Every university system in the world rests on public funding, there has never been an alternative model at any time in history. We have universities for literally the same reason that we have roads and armies.
5439621112
Reposted by Gregory Marton
Tal Haklay @talhaklay.bsky.social · 06/03/2025
1/13 LLM circuits tell us where the computation happens inside the model—but the computation varies by token position, a key detail often ignored! We propose a method to automatically find position-aware circuits, improving faithfulness while keeping circuits compact. 🧵👇
1268
Reposted by Gregory Marton
Ethan Mollick @emollick.bsky.social · 25/02/2025
This is a crazy paper. Fine-tuning a big GPT-4o on a small amount of insecure code or even "bad numbers" (like 666) makes them misaligned in almost everything else. They are more likely to start offering misinformation, spouting anti-human values, and talk about admiring dictators. Why is unclear.
721443
Gregory Marton @gregory-marton.bsky.social · 14/02/2025
Removing the gears part results in better performance, and that's surprising because it feels different from how humans learn. Perhaps relatedly, though anecdotally, telling e.g. an image generator what you didn't like about the previous response results in more, not less, of what you didn't like.
040
Reposted by Gregory Marton
Joseph Howley @illdottore.bsky.social · 12/02/2025
you fucked up a perfectly good computer is what you did. look at it. it's got innumeracy
Screenshot of a twitter post showing that the latest openAI commercial model is better than previous models at doing arithmetic but still cannot reliably produce the correct answer of multiplication problems with values greater than 11 x 11.  It's supposed to be impressive I think
1083161624
Gregory Marton @gregory-marton.bsky.social · 06/02/2025
«Kim pointed to newer introductory offerings such as “Python for Humanities and Social Sciences,” “AI for Future Presidents” and “C Programming Language and Linux.”» and it's still available free online www.edx.org/cs50 Love the homage to Richard Muller, too!
yaledailynews.com
“This was CS50”: Yale ends largest computer science course
After a decade of partnership with Harvard, Yale’s CS50 course will no longer be offered starting in fall 2025 due to limited funding and an expanding computer science department.
010
Reposted by Gregory Marton
Ethan Mollick @emollick.bsky.social · 06/02/2025
Are linguists paying a lot of attention to LLMs? Because this seems like a fascinating finding with large implications: LLMs share highly abstract grammatical concept representations, even across unrelated languages, so even models trained mostly on English do well in other languages.
1116413
Reposted by Gregory Marton
Fernanda Ferreira @fernandaedi.bsky.social · 05/02/2025
Tech oligarchs made their fortunes thanks in large part to government funded research done by scientists based in universities. The tech industry’s complicity in dismantling these govt agencies and higher ed is not only immoral, it’s also shortsighted. Where will new science breakthroughs come from?
5384
Reposted by Gregory Marton
Jeff Dean @jeffdean.bsky.social · 05/02/2025
We launched a bunch of Gemini 2.0 models today. Compared to the 1.5 series models, each of the 2.0 models is generally better than the "one size up" model in the 1.5 series. 2.0 Flash & Flash-Lite set new standards in the quality/cost Pareto frontier. More details: blog.google/technology/g...
49414
Reposted by Gregory Marton
Sung Kim @sungkim.bsky.social · 04/02/2025
open-Deep-Research by huggingface as posted by @aymeric-roucher.bsky.social An entirely open agent that can: navigate the web autonomously, scroll and search through pages, download and manipulate files, run calculation on data...
1134
Reposted by Gregory Marton
Samuel @samuel.fm · 03/02/2025
man stepping on rake labelled 4o

man skatedboarding down stairs on a rake before getting hit in face labelled o3-mini-high
523818
Reposted by Gregory Marton
Sung Kim @sungkim.bsky.social · 01/02/2025
Exa & Deepseek R1 Chat App Exa & Deepseek Chat App is a free and open-source chat app that uses Exa's API for web search and Deepseek R1 LLM for reasoning. github.com/exa-labs/exa...
github.com
GitHub - exa-labs/exa-deepseek-chat: A simple open-source chat app that uses Exa's API for web search and Deepseek R1 for reasoning
A simple open-source chat app that uses Exa's API for web search and Deepseek R1 for reasoning - exa-labs/exa-deepseek-chat
1122
Reposted by Gregory Marton
Mike Dickison @adzebill.bsky.social · 01/02/2025
The Internet Archive has to date downloaded 500 terabytes of US government websites, which it crawls at the end of every presidential term. The whole archive is fully searchable. This effort's housed by a donation-funded nonprofit, not a branch of the US government. blog.archive.org/2024/05/08/e...
blog.archive.org
End of Term Web Archive – Preserving the Transition of a Nation | Internet Archive Blogs
4803276812099
Reposted by Gregory Marton
Nicole Hennig @nic221.bsky.social · 30/01/2025
Researchers claim Linux kernel tweak could reduce data center energy use by 30% www.techspot.com/news/106501-linux-… #AI #climate
Text Shot: Professor Karsten emphasized the potential global impact of this development, noting that if major tech companies like Amazon, Google, and Meta choose to implement this method in their data centers, it could lead to savings of gigawatt-hours of energy worldwide.
031
Reposted by Gregory Marton
Karen Hao @karenhao.bsky.social · 27/01/2025
As someone who has reported on AI for 7 years and covered China tech as well, I think the biggest lesson to be drawn from DeepSeek is the huge cracks it illustrates with the current dominant paradigm of AI development. A long thread. 1/
21161242339
Reposted by Gregory Marton
mr. TIM @timkellogg.me · 26/01/2025
Explainer: What's R1 and Everything Else This is an attempt to consolidate the dizzying rate of AI developments since Christmas. If you're into AI but not deep enough, this should get you oriented again. timkellogg.me/blog/2025/01...
The image depicts a monumental statue of Buddha, emphasizing serenity and grandeur. The statue's intricate design captures traditional Buddhist features, including a meditative posture with hands placed in a symbolic gesture, flowing robes, and a calm facial expression exuding peace. The perspective highlights the statue's immense size against a minimalistic white sky background, underscoring its significance as a spiritual and cultural landmark.
411526
Reposted by Gregory Marton
Boston Tom Levenson @tomlevenson.bsky.social · 23/01/2025
I'm not sure if people realize how quickly the Trumpzis can do enormous damage to US science, from basic research to translation. Really fast. REALLY fast. Labs with decades of irreplaceable domain and technique knowledge can break apart with a surprisingly short funding gap. When they're gone...1/
21940318
Reposted by Gregory Marton
Ethan Mollick @emollick.bsky.social · 24/01/2025
Next big thing for brands: knowing what sites agents prefer. If you ask for stock prices, Claude with Computer Use goes to Yahoo Finance while Operator does a Bing search Operator loves buying from the top search result on Bing. Claude has direct preferences like 1-800-Flowers We don't know why
7747
Reposted by Gregory Marton
Melanie Mitchell @melaniemitchell.bsky.social · 23/01/2025
Worth also pointing out that there are many "tests so easy no AI system can pass them". Moravec's paradox remains. E.g., arxiv.org/abs/2404.12390
arxiv.org
BLINK: Multimodal Large Language Models Can See but Not Perceive
We introduce Blink, a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations. Most of the Blink tasks can be solved by huma...
711534
Reposted by Gregory Marton
Ethan Mollick @emollick.bsky.social · 22/01/2025
The new ability of AI video creators to add real people and products to scenes with just an image is likely to increase the utility (& more worryingly, misuse) of AI video. Here I made Shakespeare at a cafe and the Girl with the Pearl Earring piloting a mech (just as Vermeer intended)
5738
Reposted by Gregory Marton
Marc Lanctot @sharky6000.bsky.social · 17/01/2025
In December, I posted about our new paper on mastering board games using internal + external planning. 👇 Here's a talk now on Youtube about it given by my awesome colleague John Schultz! www.youtube.com/watch?v=JyxE...
youtube.com
John Schultz, DeepMind, Mastering Board Games by External and Internal Planning with Language Models
YouTube video by AI4All
13511
Gregory Marton @gregory-marton.bsky.social · 17/01/2025
Explainability focuses on finding *directions* in representation space that correspond to concepts, and strong LRH posits that this may be the only kind of representation to look for. Not so, counterexample given where magnitude matters orthogonally. aclanthology.org/2024.blackbo...
aclanthology.org
Recurrent Neural Networks Learn to Store and Generate Sequences using Non-Linear Representations
Róbert Csordás, Christopher Potts, Christopher D Manning, Atticus Geiger. Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP. 2024.
010
Gregory Marton @gregory-marton.bsky.social · 16/01/2025
"Titans", as opposed to Transformers, treat attention as short-term memory and extend the possible context window by using an additional neural memory that lives just as long as in-context document ingestion and query ("test time"), and controlled by surprise and decay. arxiv.org/pdf/2501.00663
arxiv.org
000
Reposted by Gregory Marton
Nathan Lambert @natolambert.bsky.social · 14/01/2025
Qwen released a 72B process reward model (PRM) on their recent math model. A good chance it's the best PRM openly available for reasoning research. We like Qwen. buff.ly/4gQV9wt
0314
Gregory Marton @gregory-marton.bsky.social · 13/01/2025
They found it helpful to pretrain by masking only nouns, verbs, and named entities, and only one at a time, rather than a random set of tokens, for languages where data are scarce.
142
Reposted by Gregory Marton
Simon Willison @simonwillison.net · 11/01/2025
The disadvantage of writing one big review of 2024 is that individual sections get lost in the noise - this part about both the improvements and deteriorations in terms of environmental impact of LLMs probably deserved its own separate post
5636
Reposted by Gregory Marton
Sung Kim @sungkim.bsky.social · 11/01/2025
Google just released TimesFM-2.0 (Time Series Foundation Model - jax & pytorch) on Hugging Face with a significant boost in accuracy and maximum context length. It is a pretrained time-series foundation model developed by Google Research for time-series forecasting. huggingface.co/google/times...
huggingface.co
google/timesfm-2.0-500m-pytorch · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
0343
Reposted by Gregory Marton
Gary Marcus @garymarcus.bsky.social · 01/01/2025
25 AI Predictions for 2025 (and a review of my almost entirely correct predictions from 2024) open.substack.com/pub/garymarc...
open.substack.com
25 AI Predictions for 2025, from Marcus on AI
With a review of last year’s predictions
89331
Gregory Marton @gregory-marton.bsky.social · 01/01/2025
My prediction in 2010 was that we would have more autonomous cars than human driven ones on the road by 2030, and I guess we'll see, but an important take is $ would be better spent on improving public transit and infrastructure. Better driver assistance is cool too, I guess. Yay LLMs, despite hype!
000
Reposted by Gregory Marton
Shriram Krishnamurthi @shriram.bsky.social · 29/12/2024
It's very fashionable to keep criticizing LLMs as "glorified autocorrect"s. I'm curious how one explains the ability to execute this prompt beautifully as an "autocorrect". (And yes, I've used many other language systems: Duolingo, Mango, Pimsleur, etc.)
I would like to practice Italian at the A1 level. Let's pretend to have a conversation in a cafe. You ask me questions in Italian. I will answer in both English and Italian. If my Italian matches my English and is grammatically correct, continue the conversation. If it does not, correct my Italian and then continue the conversation.
16434
Reposted by Gregory Marton
Iris van Rooij 💭 @irisvanrooij.bsky.social · 28/12/2024
Published version, here: van Rooij, I., Guest, O., et al. Reclaiming AI as a Theoretical Tool for Cognitive Science. Comput Brain Behav 7, 616–636 (2024). doi.org/10.1007/s421...
doi.org
Reclaiming AI as a Theoretical Tool for Cognitive Science - Computational Brain & Behavior
The idea that human cognition is, or can be understood as, a form of computation is a useful conceptual tool for cognitive science. It was a foundational assumption during the birth of cognitive scien...
68621
Gregory Marton @gregory-marton.bsky.social · 22/12/2024
Genius! For medical device communication, do not use wireless, which is easy to snoop or jam, nor implant actual wires, ugh, but use the human body itself "as the communication medium for the devices in someone's body-area network." #IoBodies wow.
031
Reposted by Gregory Marton
Ethan Mollick @emollick.bsky.social · 21/12/2024
Basically think of the o3 results as validating Douglas Adams as the science fiction author most right about AI. When given longer to think, the AI can generate answers to very hard questions, but the cost is very high, it is hard to verify, & you have to make sure you ask the right question first.
718536
Reposted by Gregory Marton
Christine Hall @brideoflinux.bsky.social · 23/10/2024
"The Free Software Foundation announced they are pursuing freedom in machine learning while not being limited to just the software but also the training data as well": The Free Software Foundation Finally Has AI / Machine Learning Apps On Their Radar - Phoronix
buff.ly
The Free Software Foundation Finally Has AI / Machine Learning Apps On Their Radar
The Free Software Foundation announced on Tuesday they have begun work on 'freedom in machine learning applications'
011
Gregory Marton @gregory-marton.bsky.social · 19/12/2024
A lot more encoding happens than generation, because e.g. to find query-relevant documents you encode them all and look for similarities in the encoded space. Improvements in encoding are thus less visible but perhaps more impactful from sustainability and quality viewpoints.
000
Reposted by Gregory Marton
Lilian Edwards @lilianedwards.bsky.social · 19/12/2024
Dear god does this really all need to happen approximately 2 days before end of days! Ill come back to you on this one 🤣 @create-glasgow.bsky.social www.gov.uk/government/c...
gov.uk
Copyright and Artificial Intelligence
2193
Reposted by Gregory Marton
Dan Meyer @ddmeyer.bsky.social · 04/12/2024
People are right now slobbering pretty hard over this AI tutor demo over on LinkedIn. I think it's a mess—pedagogically, socially, and mathematically. What do you notice?
194110
Reposted by Gregory Marton
Benno Krojer @bennokrojer.bsky.social · 14/12/2024
Amazing line up of speakers this afternoon, glad I chose to attend @suhr.bsky.social talked about interactive language use in games, specifically their latest project on studying how people cooperate/talk in Portal 2 Tom Griffiths, among other insights, showed tasks where CoT*hurts* performance!
1101
Reposted by Gregory Marton
Laura Spencer, Ed.D. @laurakspencer.com · 12/12/2024
Used Gemini Live tonight to discuss CA’s new law AB 2013. The conversation flowed easily and could definitely help with brainstorming complex ideas. #EduSky #AIEdu
021
Reposted by Gregory Marton
Vukosi Marivate @vukosi.bsky.social · 12/12/2024
U.S. math scores drop on major international test | buff.ly/3VNRNSv
011
Reposted by Gregory Marton
Ben Schmidt @bschmidt.bsky.social · 11/12/2024
Google quietly updated their ngrams viewer again this year. The books used appear to be extremely different yet again--the rate of the words "she said" are about 60% what they were in the 20th century compared to the 2019 release, and just 20% compared to 2009. But there's a catch:
Four line charts showing the same words -- 'she said' plotted in four different corpora on the google ngrams viewer. The words are much less common in the 2019 and most recent version through the 20th century, and spike up sharply after 2000.
521579