Marin's 535B-parameter hero run was originally planned as a 360B model.
A new blog from Rafal Wojdyla, MTS, explains how a change in sharding strategy, Expert Parallelism, made room for a ~50% larger model at the same ~23B active parameters, and what it costs.
🔗 openathena.ai/blog/expert-...