Sign in

Sung Kim

@sungkim.bsky.social
8.7K followers 1.2K following 9.1K posts

A business analyst at heart who enjoys delving into AI, ML, data engineering, data science, data analytics, and modeling. My views are my own. You can also find me at threads: @sung.kim.mw

PostsRepliesMedia
Sung Kim @sungkim.bsky.social · 5h
OK, where is the Codex reset? My weekly limit is in single digit.
110
Sung Kim @sungkim.bsky.social · 6h
I may not be the only one trying to use Dots for software development because of its allegedly unlimited tokens, which don’t count against my usage. OpenAI now seems to be actively limiting Dots’ ability to develop software. I’m running into blocker after blocker.
2160
Sung Kim @sungkim.bsky.social · 16h
This is a GREAT idea for wedding videographers, except I’d use AI to generate these videos. Brides and grooms are spending thousands on film crews, scripts and costumes to star in elaborate romance dramas, a rare bright spot in a shrinking marriage market. www.bloomberg.com/news/article...
bloomberg.com
Chinese Couples Are Turning Wedding Videos Into Epic Productions
Brides and grooms are spending thousands on film crews, scripts and costumes to star in elaborate romance dramas, a rare bright spot in a shrinking marriage market.
1170
Sung Kim @sungkim.bsky.social · 17h
Maybe I can AI clone myself and let them remote work for companies??? Tavus' Griffin, the first(?) model to pass video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video.
175314
Sung Kim @sungkim.bsky.social · 18h
AI replacing accountants was predicted very early on, but it actually seems to be happening now... "Human Baselines for Benchmarks: AI Now Outperforms Junior Accountants" www.mercor.com/blog/human-b...
1140
Sung Kim @sungkim.bsky.social · 18h
My review of OpenAI’s Dots: As usual with OpenAI, it feels like yet another half-baked product. OpenAI needs new leadership that is less focused on reacting to competitors and more focused on building polished, fully realized products.
4351
Sung Kim @sungkim.bsky.social · 18h
I cannot emphasize enough that software development with OpenAI's dots is really slow. It is still working on same project.
020
Sung Kim @sungkim.bsky.social · 01/10/2026
My AI agent dev workflow: I start by writing a design document for a new feature or a refactor of an existing part of the codebase. I review and revise that document repeatedly, often around 10 times, until I can no longer identify any significant issues or ambiguities.
192
Sung Kim @sungkim.bsky.social · 01/10/2026
Scaling Pre-Training in Practice: A Hierarchical Approach by Jordan Sassoon It’ll guide you through the choices we make at different GPU counts, taking a 30B-A3B MoE model from 16 - 512 B200 GPUs with 35% MFU & near-linear scaling. aleph-alpha.com/en/blog/scal...
190
Sung Kim @sungkim.bsky.social · 01/10/2026
Learning from the 7,780 environments Xiaomi open-sourced for MiMo. In RL environments, reward design is everything huggingface.co/spaces/FineE...
0224
Sung Kim @sungkim.bsky.social · 01/10/2026
Reshaping Monte-Carlo Tree Search 2FFS A new tree search algorithm that combines multi-fidelity bandits to resolve the fundamental trade-off: should we use cheaper, approximated evaluations, or expansive but accurate samplings? Paper: arxiv.org/abs/2606.01708
0191
Sung Kim @sungkim.bsky.social · 01/10/2026
The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA! - Natively manage their own context - Treat context as a file - Learn policies in CLM weights, no harness
2554
Sung Kim @sungkim.bsky.social · 01/10/2026
Tokenization: A Survey for Modern NLP Over the past ~8 months, 32 (!) tokenizer researchers put together the comprehensive survey of the field. www.alphaxiv.org/abs/2609.tok...
0183
Sung Kim @sungkim.bsky.social · 01/10/2026
Word in the street on gpt-6.1-sol being slow....
1160
Sung Kim @sungkim.bsky.social · 01/10/2026
Christopher Degnan of Snowflake announces that he is joining Cognition Labs. Matan Grinberg of Factory AI accuses Degnan, who had served as an advisor to Factory AI, of stealing confidential information and taking it to Cognition Labs.
100
Sung Kim @sungkim.bsky.social · 30/09/2026
I really miss the old days when one of the most exciting things in machine learning was XGBoost.
4722
Sung Kim @sungkim.bsky.social · 30/09/2026
We're back to three way race again.
3212
Sung Kim @sungkim.bsky.social · 30/09/2026
Is anyone using cloud 'dev' computer like Cube Computer (cube.computer) or boat (boat.dev) for your development as well as your agents? They seem affordable.
cube.computer
Cube: an always-on cloud computer for Claude Code & Codex
Run coding agents on your Mac or a personal cloud computer. Keep projects, terminals, and previews together in Cube.
271
Sung Kim @sungkim.bsky.social · 30/09/2026
It's a bit funny that OpenAI's Tibo went from one of the most beloved people to one of the most disliked people in a day. Now, they're openly fighting. 🤣🤣🤣
020
Sung Kim @sungkim.bsky.social · 30/09/2026
I've been developing with OpenAI's Dots since yesterday, and here's my take: 1. When developing with Codex, it provides constant feedback on what it is working on. Dots does not provide any feedback unless asked.
240
Sung Kim @sungkim.bsky.social · 30/09/2026
I'm so confused. What is Codex task in this context?
310
Sung Kim @sungkim.bsky.social · 30/09/2026
In the old days, young people entering the workforce had a technological advantage because they had grown up with tools that older generations had to learn later in life. But AI is relatively new to everyone, and judging its performance requires knowledge of the business domain.
2162
Sung Kim @sungkim.bsky.social · 30/09/2026
My review of OpenAI dots (openai.com/index/introd...). I'm using it as a coding agent. It seems to be working well since it uses gpt-6-astra and its usage is NOT counted against my usage. I'll be using this a lot while it's kind of free.
openai.com
Introducing dots
Dots by OpenAI are proactive assistants that can keep working across complex projects and everyday tasks. Learn how dots help you stay in control while work moves forward.
2192
Sung Kim @sungkim.bsky.social · 30/09/2026
How about creating a decision model or Yet Anther Jev, what if someone create a LLM-based decision tree model. Just saying... Input: { "question": "Does this PR require security review?", "context": "...", "branches": [ "security_review", "normal_review", "skip_review" ] }
360
Sung Kim @sungkim.bsky.social · 29/09/2026
It feels like OpenAI developed dots (chatgpt.com/features/dots) to counter Claude tag (www.anthropic.com/news/introdu...), but made it cute to counter Meta's Muse (ai.meta.com/muse/).
chatgpt.com
Dots: Always-on agents built to handle everything | ChatGPT
Meet your dot. It learns your priorities, takes on ongoing tasks, and keeps making progress between conversations, giving you more time for what matters.
0140
Sung Kim @sungkim.bsky.social · 29/09/2026
Nvidia's Kumo Tabular, a new family of foundation models for tabular data - open weights, open-source software, and a permissive license for commercial use. HuggingFace: huggingface.co/nvidia/Kumo-... GitHub: github.com/NVIDIA/struc...
2322
Sung Kim @sungkim.bsky.social · 29/09/2026
Jev's moat lasted like a day. LiquidAI announces d1, their decision model. > Decision models page: docs.liquid.ai/lfm/models/d... > Migration guide: docs.liquid.ai/guides/decis... > Demo tutorial: github.com/Liquid4All/c...
291
Sung Kim @sungkim.bsky.social · 29/09/2026
OpenAI dots is FREE for a limited time. Dots uses gpt-6-astra I can create unlimited dots. I'm just using dots for all my software development.
2141
Sung Kim @sungkim.bsky.social · 29/09/2026
OpenAI: we're aligning our API pricing with subscription pricing. Now, you can just use your subscription with third-party apps. Thanks, I guess...
1210
Sung Kim @sungkim.bsky.social · 29/09/2026
claude tag meta muse openai dot
020
Sung Kim @sungkim.bsky.social · 29/09/2026
There probably will be a codex reset tomorrow, so you may want to drain your usage.
3160
Sung Kim @sungkim.bsky.social · 29/09/2026
When you feel you need a head-inducing list of configurable ACL for your shell/VM, to run your AI agents. github.com/NVIDIA/OpenS...
github.com
GitHub - NVIDIA/OpenShell: OpenShell is the safe, private runtime for autonomous AI agents.
OpenShell is the safe, private runtime for autonomous AI agents. - NVIDIA/OpenShell
0212
Sung Kim @sungkim.bsky.social · 29/09/2026
Is your GitHub Actions' cost hitting 3 or 4 figures? May I recommend self-hosting your own Git Runners?
180
Sung Kim @sungkim.bsky.social · 29/09/2026
My thoughts on Instinct AI: It's too cool for its own good and it does not support Android's SMS/RCS.
010
Sung Kim @sungkim.bsky.social · 28/09/2026
Google's RRSI: Regularized Recursive Self-Improvement of Agent Harnesses github.com/google-resea...
0232
Sung Kim @sungkim.bsky.social · 28/09/2026
I don’t know about you, but I trust Meta more than some startup with a pretty UI wrapped around an AI agent. Time and time again, we find out that these products launched with serious security gaps that should have been addressed from day one. Wajo wajo.ai/join-wajo
wajo.ai
Join Wajo
Give Fo the things you need done. Fo calls, emails, books and pays – then follows through until the job is finished.
340
Sung Kim @sungkim.bsky.social · 28/09/2026
Interesting...
092
Reposted by Sung Kim
mr. TIM @timkellogg.me · 28/09/2026
Claude 5.5 Sonnet is live and it’s roughly Opus 5.5 but cheaper www.anthropic.com/claude-sonne...
A benchmark comparison table titled "Claude Sonnet 5.5" comparing four models: Sonnet 5.5, Sonnet 5, Opus 5.5, and GPT-6 Sol.
 * Agentic coding (Terminal-Bench 4.0): Sonnet 5.5: 70.6%; Sonnet 5: 10.3%; Opus 5.5: 66.4%; GPT-6 Sol: N/A.
 * Agentic coding (FrontierCode 1.1 Main): Sonnet 5.5: 46.2% (Max) / 52.1% (Xhigh); Sonnet 5: 42.4%; Opus 5.5: 54.4%; GPT-6 Sol: 49.3%.
 * Agentic coding (CursorBench 4.0): Sonnet 5.5: 55.5%; Sonnet 5: 34.1%; Opus 5.5: 57.8%; GPT-6 Sol: N/A.
 * Knowledge work (GDPval-AA v2.1): Sonnet 5.5: 1844; Sonnet 5: 1449; Opus 5.5: 1846; GPT-6 Sol: 1487.
 * Knowledge work (AA-Briefcase v1.1): Sonnet 5.5: 1811; Sonnet 5: 1359; Opus 5.5: 1822; GPT-6 Sol: 1483.
 * Multidisciplinary reasoning (Humanity's Last Exam with tools): Sonnet 5.5: 64.5%; Sonnet 5: 54.9%; Opus 5.5: 67.7%; GPT-6 Sol: N/A.
 * Computer use (OSWorld 2.1 partial): Sonnet 5.5: 80.1%; Sonnet 5: 57.0%; Opus 5.5: 81.8%; GPT-6 Sol: N/A.
 * Visual chart recognition (Chartography no tools): Sonnet 5.5: 61.6%; Sonnet 5: 15.6%; Opus 5.5: 64.4%; GPT-6 Sol: 53.6%.
Footnotes provide methodology notes regarding evaluation settings, Artificial Analysis pre-release testing details, and recent bug fixes affecting GPT-6 Sol benchmark scores.
A line graph titled "Agentic coding by effort level" on the CursorBench 4.0 benchmark, plotting Score (%) on the linear y-axis (20% to 60%) against Cost per task in USD on a logarithmic x-axis ($0.5 to $10+).
The chart compares four models:
 * Sonnet 5.5 (blue line with labeled effort levels): Starts at "Low" (~$0.50, 36%), moving through "Med" ($0.70, 39%), "High" ($1.70, 48%), "Xhigh" ($3.80, 53%), and "Max" ($9.50, ~55.5%).
 * Opus 5.5 (orange line): Tracks closely above Sonnet 5.5 at higher cost points, spanning from ~$1.20 per task (~44%) up to ~$12.50 per task (~58%).
 * Sonnet 5 (grey line): Shows lower accuracy relative to cost, ranging from ~$0.90 per task (~25%) to ~$8.00 per task (~42%).
 * GPT-5.6 Sol (light green line): Represents the lowest trajectory, ranging from ~$1.40 per task (~24%) to ~$7.00 per task (~34%).
Footnote: "CursorBench 4.0 does not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here."
2020623
Sung Kim @sungkim.bsky.social · 28/09/2026
Manus 2.0 Can someone do a comparison between Meta's Muse and Manus 2.0? manus.im/blog/introdu...
manus.im
Introducing Manus 2.0
Manus 2.0 is here, with the new Cascade agent harness, Cloud Computer and Automations, Manus Studio for professional creation, and Cue, a new app for personal agents.
190
Sung Kim @sungkim.bsky.social · 28/09/2026
From my estimation, gpt-6-astra runs for about 2 days on a 20x plan and claude-opus-5.5 runs for about 2.5 days on a 20x plan. I do not hit 5 hour limit on claude-opus-5.5.
2130
Sung Kim @sungkim.bsky.social · 27/09/2026
So... I’ve automated the code review process for my personal projects, and it’s actually getting kind of boring now that agentic AI coding harnesses can run for days at a time.
360
Sung Kim @sungkim.bsky.social · 27/09/2026
MiniMax kind of launched M3.1-Flash-Preview, only for MiniMax Code users.
0120
Sung Kim @sungkim.bsky.social · 27/09/2026
If you're buying an iPhone 18 Pro in the U.S., consider the iPhone 18 Pro Max instead of the regular Pro. It's all about the cellular modem. The U.S. iPhone 18 Pro Max uses Qualcomm's modem, while the iPhone 18 Pro uses Apple's new C2 modem.
5141
Sung Kim @sungkim.bsky.social · 27/09/2026
It seems Tokyo Central Market is expanding in the Los Angeles area. The funny thing is that it’s owned by the same company that owns Don Quijote in Japan.
190
Sung Kim @sungkim.bsky.social · 26/09/2026
Yet Another Jev Variant. I should create an acronym - YAJV! Julia 1: A decision model that runs on almost anything Blog: supersoniclabs.ia.br/julia-1/ Model: huggingface.co/SupersonicLa...
supersoniclabs.ia.br
Introducing Julia 1 | Supersonic Labs
Julia 1 is our compact decision model. Explore the research, evaluation results, limitations, and model repository.
4938
Sung Kim @sungkim.bsky.social · 26/09/2026
I got a codex reset today. Thanks! I guess. Only like 10 hours into my regularly scheduled weekly reset.
2120
Sung Kim @sungkim.bsky.social · 26/09/2026
How can I tell if a model is smart? See if it understands sarcasm. Claude: Yes GPT-6: Yes Muse Spark: No Devin SWE2: No Cursor Composer: No
2130
Sung Kim @sungkim.bsky.social · 25/09/2026
State of AI coding agents as I see it, based on my experience as of September 25, 2026: - Claude Code: I use the CLI version, and Claude Opus 5.5 is working well. - Codex: I use the GUI version, and GPT-6 Astra is working well.
6300
Sung Kim @sungkim.bsky.social · 25/09/2026
I'm still evaluating, but gpt-6-astra is better at a long-running agentic implementation task than claude-opus-5.5.
1140
Sung Kim @sungkim.bsky.social · 24/09/2026
Is it just me or your social feed is also all muse and dhh?
560