opliko @opliko.dev · 07/10/2026Their CEO said they'd be launching a model (presumably this one) in summer back in July... Also, it's not even really done training - RL is ongoing, hence the preview, and final release w/ weights is slated for end of October. 040
opliko @opliko.dev · 05/10/2026Same experience, worked at every university I tried it at... Except the one I attend and got the credentials from... 030
opliko @opliko.dev · 03/10/2026Heck, coming out to say "we gave it our shot, beat Mistral on our ~first release, but are still more behind even smaller Chinese OSS models than China is behind the frontier" *should* serve as a wake up call that EU needs more AI labs competing, because it's dire out here. 020
opliko @opliko.dev · 03/10/2026The annoying thing is: it's good to have an EU competition, even if just on the smaller model front. They did beat Mistral Small 4, which is their newest model in this segment. But they should just be honest about this, and not try to create a fake reality where EU is at any frontier here 160
opliko @opliko.dev · 03/10/2026(And yes, AA is not ideal and all benchmarks are not *great*, but at some degree of difference across different benches it can't really be just benchmaxxing. And they've used many of the same benches for their comparisons) Also correction: Nemotron is slightly behind them on their benches overall 140
opliko @opliko.dev · 03/10/2026The best model they compare against is Nemotron 3 Super, which beats them on benchmarks slightly. Wonder why not compare to Qwen3.8 Flash Next which is even lighter in active parameters, being a 125B/6B MoE (+engrams)... Ah, that's why: 150
opliko @opliko.dev · 02/10/2026I could see ECB possibly working, since it does leak plaintext patterns. Probably very badly, but still - has a chance at least. Any cipher mode with feedback, chaining or even counter should make it impossible though. 110
opliko @opliko.dev · 02/10/2026And openmdw.ai seems to be slowly gaining momentum tooopenmdw.aiOpenMDWA permissive license crafted for machine-learning models 100
opliko @opliko.dev · 02/10/2026So overall open weight is more accurate not because they're not open, but because they're not the same kind of software. Just like open data. And yes, some are trying more limiting custom licenses, and even if they're unenforceable it sucks, but there are a lot of models released under MIT/Apache 101
opliko @opliko.dev · 02/10/2026But there is a part of open LLMs that has a source code: the actual runner that takes the weights and lets you do inference on them. And that part is FOSS for basically all open weight models. VLLM, SGLang and transformers are all Apache 2.0. 110
opliko @opliko.dev · 02/10/2026We don't call FOSS software written in JetBrains IDEs or Visual Studio not FOSS, let's be serious, the tool used to create a software artifact can be anything and it'll still be open. Weights are not binaries compiled from some source code, but, again, closer to databases. 120
opliko @opliko.dev · 02/10/2026Unless you mean only weights and implementations (typically in engines line vllm or sglang) being open, and not the entire training dataset and environment. It's good some companies are releasing more of that too (NVIDIA, ai2, now Xiaomi releasing RL env), but it's not a requirement to be open IMO 110
opliko @opliko.dev · 02/10/2026At least in the EU, the weights are most likely non-copyrightable and not protected unless the company releasing them is from the EU (sui generis database protection is only for EU nationals and residents, and even then it's only valid for 15 years not the copyrights practically-forever) 110
opliko @opliko.dev · 01/10/2026Gemini Oganesson as a jev-like fast decision model (p50 time to answer at 0.7ms) 010
opliko @opliko.dev · 01/10/2026And, well, it definitely doesn't make LLMs any more accurate or correct than when they're sampling more randomly, if anything the effect on output quality is typically negative overall. 000
opliko @opliko.dev · 01/10/2026Case in point: it's possible to make LLMs deterministic too, it's just rarely useful and would result in slower inference at scale in practice (technically making the model deterministic is trivial, but a buch of optimizations add variance when batching requests, which is required for high perf) 210
opliko @opliko.dev · 01/10/2026But, like, explaining LLMs has clearly become a part of what she does professionally. It's obviously not the only part and I get that simplifying an academic career to explaining one idea that's not its entire focus is oversimplifying, but it being only a part of it doesn't change the point really. 030
opliko @opliko.dev · 01/10/2026(I mean, most of her classes are more fundamental NLP and computationally linguistics AFAICT, but at least some seminars she was an instructor for included topics around LLMs. And even if we assume her teaching isn't LLM-related her publications still are) 150
opliko @opliko.dev · 01/10/2026That's some harsh criticism of her career and publications. I know neither university professors nor most authors are paid that well, but I'd imagine it's enough to be considered professional. She's not paid to explain LLMs on Bluesky. She is paid to explain LLMs in her books, papers and classes. 140
opliko @opliko.dev · 28/09/2026My intuition would be that it depends on # of experts? Mistral has only 128 experts with 4 active at one time, meanwhile gemma 26b also has 128, but with 8 active (despite much smaller size), ds4f has 256 with 6 active and Qwen3.8f does 512/10 So each one is probably more "fungible" in these models 000
opliko @opliko.dev · 26/09/2026openai.com/index/scalin... It's incredibly clear they've been working *around* the original architecture just to scale as ChatGPT exploded in popularity, and it *is* impressive they managed toopenai.comScaling PostgreSQL to power 800 million ChatGPT usersAn inside look at how OpenAI scaled PostgreSQL to millions of queries per second using replicas, caching, rate limiting, and workload isolation. 060
opliko @opliko.dev · 26/09/2026They had a very telling technical blog post about scaling PostgreSQL that's clearly trying to be some success story, but to everyone with any relevant experience was essentially a legacy architecture horror story, because the only way you end up doing things they did is by not planning for scale. 150
Reposted by oplikoTeefa @tifaret.bsky.social · 24/09/2026Introducing ONTOPAUSE™: An AI solution that automatically generates, resolves, and hides important ontological clarifications about the language you’re using, so everyone can reaffirm exactly what they don’t mean before going back to talking normally. Start not-stopping today! 110017
opliko @opliko.dev · 25/09/2026Also, this was mostly about pre-train data, AFAICT mid-train (things like context extension, instruction following, reasoning examples) and post training use more synthetic data (or largely RL for the latter, which is synthetic I guess but not really training data in the same way) 020
opliko @opliko.dev · 25/09/2026Also, Nemotron was trying to be an actually open model with at least most of the data mix released as well, most open weight models don't release similar data or almost any details on their data sources really. At most they give total token numbers by category of data (code, general knowledge, etc.) 120
opliko @opliko.dev · 25/09/2026Nemotron 3 Nano claims ~33% synthetic, though I'd say an interesting issue nowadays would be defining what even is synthetic, because a lot more of the data was processed by LLMs in some way (from simple cleanup to rephrasing). With some napkin math I think you'd get above 50% by counting generously 120
Reposted by oplikoSamuel @samuel.fm · 24/09/2026LLMs will never beat human developers at RSI. they don't even have wrists 41074
opliko @opliko.dev · 23/09/2026(It was a short moment where Spark prices jumped and R9700 were at a good price for some reason. They also jumped now though. Now you can only get two and be like third of a way to the next one for the price of one spark :V) 000
opliko @opliko.dev · 23/09/2026There literally was a moment where at least in Poland you could almost get 4xR9700 for the price of a single GB10/DGX Spark (tho you'd need the rest of a PC if you didn't have something that could fit them) 100
opliko @opliko.dev · 23/09/2026From recent work it also seems like ~30B is a good size for a very fast jev-like model (with e.g. DiffusionGemma with structured output being one of the best options right now), but dedicating a 32GB GPU for that is probably not even the best use of resources for people with only one such card :) 010
opliko @opliko.dev · 23/09/2026I mean, R9700 was even pretty recently well below half the price of a 128gb FW desktop and it offers much faster compute with better low quant support and over 2x memory bandwidth, but yeah: with ~180-300B being the current sweet spot for models a 128GB mini PC is by far the best consumer option. 210
opliko @opliko.dev · 23/09/2026Sadly Jensen was totally right that the more you buy the more you save :V 100
opliko @opliko.dev · 23/09/2026And even then ROI won't be very quick vs properly hosted services and you need someone to manage it, even if not as a full time job at this scale. So if you also value their time doing that properly it may still end be hard to justify financially. 100
opliko @opliko.dev · 23/09/2026It starts becoming more sensible when you have multiple user and are getting to at least HPC workstation levels (so ~100k USD). And even then that's conditional on how much utilization you can get and if you're fine with the "mid sized" models. 100
opliko @opliko.dev · 23/09/2026I mean, yeah: economies of scale are brutal and unless you have a separate reason to host a model the HW is not worth it at consumer level, because you can always get higher quality cheaper. A year is just a long time in LLMs now and models are getting better! 100
opliko @opliko.dev · 23/09/2026Depends on the workflow tbh, IME 3.6-3.8 are actually good and not just benchmaxxed, better than last years frontier, but I will concur it may be due to the specific workflows I've used them in (around cybersecurity) have been a training focus only more recently. 100
opliko @opliko.dev · 23/09/2026Also, sure, small models doing more tests time compute does mean they're even slower on consumer hardware, but we are also getting a lot of optimizations that should be helpful, from smaller per-token kv size, through better 4-bit quants, often "from factory" now, to far better offload with engrams 010
opliko @opliko.dev · 23/09/2026But if you need to self host for some reason, or are totally fine with being around a year behind frontier? It's pretty absurd how good the small models have gotten and that they might still be getting better (we'll see with Qwen 4 27b tho) 110
opliko @opliko.dev · 23/09/2026Now, does this mean it's "worth it" to host local models? If you don't have a specific need probably not. It would take a pretty long time to make back the investment compared to API prcinng of hosted smaller models (like dsv4f or luna) even if you considered their tokens equivalent (they're not) 110
opliko @opliko.dev · 23/09/2026Sure, there are and will be limits to small local models. It's worse than last years frontier on knowledge for one. So it's not like it just wins on everything vs old frontier. But it'd still be competing with them if it was released last year. Potentially running on a single consumer GPU. 110
opliko @opliko.dev · 23/09/2026Most benchmarks seem to place it around Opus 4.6 territory, even if we assume benchmarks overstate its abilities, it's probably still at least ~Opus 4.5 level in many tasks. That's a flagship model from less than a year ago, that was $5/$25 then. Potentially running on a single consumer GPU. 220
opliko @opliko.dev · 22/09/2026(Ignore plagiarism detection, Pangram found this paper plagiarized itself) I'm also not sure how much I'd trust the authors pilot results considering neither of them have any law related credentials or experience AFAICT. But neither do I, hence this meta-criticism rather than something merit-based 130
opliko @opliko.dev · 22/09/2026It feels reasonable to me, though I wish the authors would at least redact their LLM writing a bit better because it smells very claudish (tho it's not 4.8 claudish yet, so I'm not confident it's not gpt) and Pangram seems to agree www.pangram.com/history/ccc5...pangram.com2609.17546v1.pdfDoes "2609.17546v1.pdf" contain AI-generated text? Pangram finds that 91% of this document is AI-generated. 121
opliko @opliko.dev · 22/09/2026Which can be true and yet misleading. A system that can find a cure for 50% of stage 3 cancer cases would revolutionize medicine. Conversly, if a blood test got Rh antibodies right just over 50% of the time it'd be worse than a constant value. 140
opliko @opliko.dev · 22/09/2026Just looking at a score out of 100% without context, as much as LLM vendors would love if everyone did just that, is not a good way to judge if an entire class of tasks is actually solved. And I'm not sure if there is any actually good solution for this sadly. 240
opliko @opliko.dev · 22/09/2026It's not an issue unique to legal profession. Every single writing quality benchmark I've seen has been clearly bullshit. And even a lot of coding benchmarks have a ton of issues, mostly being useful for comparisons between models than some overall assessment of their capabilities vs humans. 140
opliko @opliko.dev · 22/09/2026I'm not qualified to judge the details of such assesments, just to say that looking at their previous results doesn't support the assertion that LLMs are bad at this, suggesting either that they're not or something was wrong with the assesment. 130
opliko @opliko.dev · 22/09/2026I don't think using bad data to say something it doesn't really say is better than admitting there isn't good data on a topic. I'm not familiar with the area, so I don't know if there is some better evaluation from the industry. Val's seems to have worked with the industry and got these results. 251
opliko @opliko.dev · 22/09/2026Like, the claim that mid-2025 general LLMs beat humans at such a task seems wrong to me even without having any experience in that field. Clearly something with the methodology was wrong, and the new version not even having human baseline doesn't fill me with confidence they fixed it. 030