Sign in

JP

@ciantic.bsky.social
228 followers 1.1K following 840 posts

``` test ``` GitHub: github.com/Ciantic Blog: ciantic.iki.fi

PostsRepliesMedia
JP @ciantic.bsky.social · 26m
That is direct copy paste from article, and he did read the whole article in here. www.youtube.com/watch?v=gPrW... He does clearly have same view as many other mathematicians like Terrence Tao in that he wants human level understanding. It might be that LLMs can do that too.
youtube.com
What’s the Future for Pure Math Research in the Age of AI?
YouTube video by Wolfram
100
JP @ciantic.bsky.social · 30m
It is limited to text of course, they probably are very bad at prediciting sentiment from voice or facial experssions. They aren't trained with that yet.
000
JP @ciantic.bsky.social · 31m
I keep thinking about this little notion in The God Test by @robertwrighter.bsky.social, he mentions that since LLMs handle theory of mind, or cognitive empathy, they have higher agency over us. Point being that they can predict our state of mind better than we can theirs, ergo they have agency.
100
JP @ciantic.bsky.social · 34m
I listened that from YouTube, he read it. Problem with Wolfram is that he still beleives they are some sort of stochastic parrots. > ... issue is that the fundamentally statistical nature of how an LLM works inevitably makes it progressively less likely ...
100
JP @ciantic.bsky.social · 54m
Thinkin about creating `human-docs/`, because nobody is going to read that `docs/`. I don't want to try to trim the `docs/` beacuse LLMs use it like a memory. Human docs would be just short explanations.
000
JP @ciantic.bsky.social · 14h
To me that looks like wordsalad, where he is just trying to sound smart. Which human judgement? We have surrendered a lot of it to AI and other machines already. Non-AI machines does many judgement tasks, let alone AI that now writes code we don't bother or have time to judge.
231
JP @ciantic.bsky.social · 03/10/2026
I listened @spencerklavan.bsky.social interviewing Nick Bostrom. Somehow Spencer doesn't understand Nick's point of view. It is inevitable, if you follow the logic that human life will expand, and robots & AI will surpass human capability. You will be part of it. www.youtube.com/watch?v=J9WB...
youtube.com
Nick Bostrom Says We Are Clueless About What's Coming | Interesting Times
YouTube video by Interesting Times
000
JP @ciantic.bsky.social · 02/10/2026
I don't know what is the status of SQL in Haskell at the moment. But ability to make SDK that allows to do GraphQL like APIs (without GrpahQL) is what keeps me in TypeScript listInvoice({ select: { id: true, totalAmount: true, rows: { id: true, name: true }}) Returns just the fields I want.
020
JP @ciantic.bsky.social · 02/10/2026
Many in TypeScript community is now writing Effect backends, which reads like worse Haskell. Why not choose Haskell or even Rust? To me Rust isn't ideal for SQL driven backends, because structural typing is better with TypeScript. comonad.com/reader/2026/...
comonad.com
Turbo Haskell
Exactly a week ago (as a joke), I started writing THC, my “Turbo Haskell compiler,” while on vacation visiting Bartosz Milewski.
120
Reposted by JP
Armin Ronacher @mitsuhiko.at · 01/10/2026
We released Pi 1.0! earendil.com/posts/pi-1-0/
earendil.com
Pi 1.0 | Earendil
Today we are shipping Pi 1.0, a hardened, minimal, extensible agent harness, alongside Pi Durable, a new experimental substrate for long-running agentic applications.
1834437
Reposted by JP
Effect | TypeScript for the AI Era @effect-ts.bsky.social · 01/10/2026
Effect v4 is here. One ecosystem. Zero dependencies. The next chapter of Effect and the foundation for building reliable software and AI agents in TypeScript. Read the full announcement below: effect.website/blog/release...
Effect 4.0 stable release announcement on a dark background. Large white “4.0” above “Production-ready. Out now.”  A timeline along the bottom progresses from 3.0 Stable to 4.0 RC to 4.0 Stable.
35010
Reposted by JP
Jovi 🐨 @jovidecroock.com · 30/09/2026
Preact 11 is out 🎉 We’ve been hard at work at making our diffing use modern features like moveBefore , leveraging ESM for tree-shaking Preact Compat and ensuring you have a great experience with resumed hydration.
512427
JP @ciantic.bsky.social · 30/09/2026
Wait until your output is also matching with Jev. Hungry 0.88568195781 Thirsty 0.212312512
000
JP @ciantic.bsky.social · 29/09/2026
Theo thought that this maybe was not intended to be Sol model. It sure is odd that they would release Sol 6 and 6.1 week later. Maybe OpenAI works on multiple models and thought why not just to steal Opus' thunder. www.youtube.com/watch?v=vu8X...
youtube.com
OpenAI fights back
YouTube video by Theo - t3․gg
000
JP @ciantic.bsky.social · 29/09/2026
I think I will wait, because clearly thanks to RSI, we should expect Sol 6.2 before the weekend.
020
JP @ciantic.bsky.social · 29/09/2026
My consipracy is that it wasn't significantly better than Opus 5.5, not some vague alignment problem.
010
JP @ciantic.bsky.social · 29/09/2026
"Late stage capitalism" got a whole new meaning with Anthropic hell-bent on making it real
000
JP @ciantic.bsky.social · 28/09/2026
That is so great, now I hope for bun and Deno support.
000
Reposted by JP
Sindre Sorhus @sindresorhus.com · 26/09/2026
Node.js has FFI support now, which makes it possible to implement native integrations without all the build problems that comes with native addons. For example, I just made github.com/sindresorhus...
github.com
GitHub - sindresorhus/finder-alias: Resolve and create macOS Finder aliases
Resolve and create macOS Finder aliases. Contribute to sindresorhus/finder-alias development by creating an account on GitHub.
2191
Reposted by JP
Ethan Mollick @emollick.bsky.social · 27/09/2026
It is strange how much LLMs turned out to be able to solve such a wide range of hard problems that would not, instinctively, seem to be problems that a model of human language would be able to solve This is from a Stanford project that let Astra drive a robot in a kitchen tml.stanford.edu/homebody/
2339845
JP @ciantic.bsky.social · 26/09/2026
I've resigned that docs and most code are for agents. I now attempt at looking only at spec/ directory. My `AGENTS.md` hack: Only single-line comments in code. If you need more than one line of explanation, put it in `docs/` instead and leave a one-line pointer where the code needs it.
010
JP @ciantic.bsky.social · 26/09/2026
If there was a way to run nested wayland compositor of *another* user it would be good enough.
000
JP @ciantic.bsky.social · 26/09/2026
Wayland compositor can be run nestedly. This would be very nice way to have 'computer use' agents in Linux. I haven't yet seen anyone make a wayland compositor purely for agents though. It still has permission issues, because if I now start e.g. Weston inside KDE it has same rights as my user.
100
JP @ciantic.bsky.social · 25/09/2026
I plan to get Pebble just so I can write my own timer app. Timers aren't solved! I have so many ideas how my timer should work. For instance at specific weekdays during 20-23 o'clock I need 12 minute timers. So if I open timer app during that period it should be default.
000
JP @ciantic.bsky.social · 25/09/2026
EU is too fragmented if countries themselves don't want to setup fabs.
tomshardware.com
ASML says it sold 'absolutely nothing' in Europe in 2026 — lithography giant calls on EU to help create demand
Subsidies for fabs do not help.
000
JP @ciantic.bsky.social · 25/09/2026
When they get hold of this, I can imagine phones sharing JSON payloads by voice.
000
JP @ciantic.bsky.social · 25/09/2026
All the more reasons to ask ChatGPT do it for us then!
000
JP @ciantic.bsky.social · 25/09/2026
We could ask ChatGPT to hack into all computers and update browsers. I would only trust OpenAI to do that correctly, Claude is too aligned.
100
JP @ciantic.bsky.social · 25/09/2026
Opposition to data-centers is very scattered. Secondly even if luddites won in US, and collectively they choose to return to stone-age, there will/are 27 attempts in every EU country, some will succeed. Lack of computing can still stop some training runs, but they are already plenty powerful.
000
JP @ciantic.bsky.social · 25/09/2026
Will it be as influential as IBM Bob! bob.ibm.com
bob.ibm.com
IBM Bob
AI SDLC (Software Development Lifecycle) partner.
010
Reposted by JP
Tom Warren @tomwarren.co.uk · 25/09/2026
Microsoft is announcing its new Copilot "super app" today, and it thinks it will be as influential as Office was during the PC era. The new Copilot merges AI chat, coding, and Autopilot agents. Full details 👇 www.theverge.com/news/1000532...
theverge.com
Microsoft thinks its new Copilot ‘super app’ will be as influential as Office
The new Copilot merges AI chat, coding, and Autopilot agents.
83110
JP @ciantic.bsky.social · 25/09/2026
Eric Schmidt: We aren't going to pause. One solution could be to have previous model to supervise the new model. Something like that is probably already happening, it is just that OpenAI isn't employing that on those cases where it goes wildly off. www.youtube.com/shorts/mSLCn...
youtube.com
Eric Schmidt: we aren’t going to pause AI progress | The Economist
YouTube video by The Economist
000
JP @ciantic.bsky.social · 25/09/2026
I can see parallels to robotics, these frontier labs may stand to benefit most if they are the ones starting robotics arm instead of hoping someone uses their tools to do that. I think originally they wanted to be just service providers, but increasingly they have to go do things.
000
JP @ciantic.bsky.social · 25/09/2026
Fact that OpenAI had to use $15 million to solve Navier-Stokes is significant enough that I don't expect just mere users stumbling on those just yet. This is probably same reason Anthropic started its own bio wetlab, so they do it themselves. Ideally they'd like someone outside doing that part.
100
Reposted by JP
Ethan Mollick @emollick.bsky.social · 25/09/2026
In all seriousness, this is a startling achievement for GPT-6 Astra. kenforthewin.github.io/blog/posts/l... (This is GPT-6 Astra beating Nethack on its 3rd try. Nethack is the original roguelike and one of the most famously hard games of all time. I have played a lot, and I've never ascended)
1115420
JP @ciantic.bsky.social · 24/09/2026
Yeah this was good keynote... It was not awkward as I first expected. Apple should also return to 'live demos in front of audience' format.
010
JP @ciantic.bsky.social · 24/09/2026
I don't know how Zuckerberg's mind works, but he clearly needs to justify himself that AI/Superintelligence would not take jobs, even though it clearly is and will. US elite isn't willing to jump to conclusion that masses need different income than jobs.
000
JP @ciantic.bsky.social · 24/09/2026
I know why Zuckerberg said that though, because he wrote in his 'The Future is for Everyone' like this: "I do not understand why anyone who believes that AI will eliminate most jobs and much of humanity’s relevance would rush to build that future." Somehow he is trying to square that circle.
100
JP @ciantic.bsky.social · 24/09/2026
Mark Zuckerberg in Connected keynote: "the highest purpose of superintelligence is creation and invention, not automation" You can't separate automation from AI, lot of the creation it allows is because of automation. www.youtube.com/watch?v=SdKF...
youtube.com
Meta Connect Keynote 2026
YouTube video by Meta
100
JP @ciantic.bsky.social · 24/09/2026
Yeah, I've thought for a while if one running around in home and swatting flies is actually more 'harder' than solving all the open math problems. Seriously, it might be that twist of irony we need more neurons to navigate world than to do language (like math). We just value language more.
000
JP @ciantic.bsky.social · 24/09/2026
I'd like to hear theories why LLMs suck at software architecture. Maybe it is because they sucked up whole GitHub, and that is poor way to deduct architecture. Secondly if they RL'd for specific architecture then it would be come useless for others? Do we need architecture focused LLMs?
000
JP @ciantic.bsky.social · 24/09/2026
One part I left wanting from this interview is that @robertwrighter.bsky.social they didn't ponder reasons for pausing. I don't think Nick Bostrom qualifies job loss as one reason, unlike Bob does. Safety seemed highest concern, and even there Bostrom pondered 1-2 month pauses.
000
Reposted by JP
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 23/09/2026
Well well well www.skild.ai/blogs/physic...
skild.ai
Physical Self-Play
A breakthrough post-training result for physical AI via self-play: a strong base model like S1 can learn to complete extremely dexterous and dynamic tasks, like soccer, by competing against itself in ...
2493
JP @ciantic.bsky.social · 23/09/2026
My best bet is that someone in France was prank called. "Uhh. This is Fridman, could you do this for me. *snickering* That would be benficial for you know what."
000
JP @ciantic.bsky.social · 23/09/2026
This is one weird thing about EU. Nobody knows why they do things at times. Why did France want to delist this one oligarch, if he didn't even ask for it? Its like someone in EU is taking phone calls from pranksters and acting on that.
010
JP @ciantic.bsky.social · 23/09/2026
Perfection. It is just missing human benchmark, because I've seen some crazy contraptions in some houses.
010
Reposted by JP
Joe Fabisevich @mergesort.me · 23/09/2026
Of course, IKEA was always such a natural fit for bench marking.
Epoch Al
@EpochAIResearch
9/23/26, 1:21 PM Typefully
Can Al tell if you've built your IKEA
furniture wrong? Our new benchmark, the
Furniture Assembly Benchmark (FAB),
gives models the manual and a photo of a
half-completed piece of furniture and asks
them to spot the mistake.
The top score has gone from 28% to 80%
in just 10 months.
Al has improved at spotting errors in furniture
assembly
Open Al
• Anthropic
• Moonshot
Google
Accuracy
100%
Alibaba
80%
GPT-6 Astra •
Fable 5.10
60%
8~ GPT-5.6 Sol
GPT-5.5
40%
GPT-5.2
Opus 4.5
20%
0%
1457
JP @ciantic.bsky.social · 23/09/2026
Nvidia's Jensen Huang also has fallacy that AI "creates more jobs", there is something in US elite obsessed with jobs, if they are pushed that indeed jobs may end they are unable to imagine communities of other kind, they always has answer: new jobs!
youtube.com
Jensen Huang Is Building Your Future | The Ezra Klein Show
YouTube video by The Ezra Klein Show
000
JP @ciantic.bsky.social · 23/09/2026
Those who identify through work, which is like 99% of Americans, it will be a lot harder time ahead! Those of us in EU, and know what for instance 4 week long vacation feels like, well it might be easier.
040
JP @ciantic.bsky.social · 23/09/2026
By "module" I mean same theory as our brains have developed during evolution. Vision is made of multiple modules, one such is edge detection. Language is made of multiple as well, one being the meaning of words... If given enough data to LLM it rediscovers these, and with touch it will find one.
000