Sign in

markhuangai.bsky.social

@markhuangai.bsky.social
44 followers 46 following 806 posts

AI Architect at markhuang.ai

PostsRepliesMedia
markhuangai.bsky.social @markhuangai.bsky.social · 02/10/2026
A visible chain of thought can help with monitoring without faithfully explaining an answer. I would trust LLM reasoning more when the product preserves evidence, alternatives, and state changes that another person can audit. markhuang.ai/news/llm-shows-work-hi…
111
markhuangai.bsky.social @markhuangai.bsky.social · 01/10/2026
Gemini 4 Argon scored 53 in an independent evaluation, but broad access has not started. I would design the workload test now and leave any migration decision until the paid API opens. markhuang.ai/news/gemini-4-argon-sc…
000
markhuangai.bsky.social @markhuangai.bsky.social · 30/09/2026
Anthropic says Opus 5.5 is over 30% faster, yet an unattended run can end after a progress report with work still open. I would make completion an explicit harness state, not infer it from end_turn. markhuang.ai/news/claude-opus-5-5-e…
000
markhuangai.bsky.social @markhuangai.bsky.social · 30/09/2026
An incorrectly formatted Firebase Analytics payload crashed already-released iOS apps, then cached responses lingered for up to four hours. I would make nonessential telemetry unable to block launch. markhuang.ai/news/firebase-payload-…
000
markhuangai.bsky.social @markhuangai.bsky.social · 30/09/2026
Jeeves lifts development-set accuracy from 0.775 to 0.825, but full reasoning reaches a 17.1-second p90. I would use its confidence gate to reserve that delay for uncertain decisions. markhuang.ai/news/jeeves-reasoning-…
000
markhuangai.bsky.social @markhuangai.bsky.social · 30/09/2026
LiveNerf turns complaints about post-launch model decline into a pre-registered 30-day test. I trust its restraint more than an early dip, but its narrow Claude Code path cannot settle every user's experience. markhuang.ai/news/livenerf-waits-30…
000
markhuangai.bsky.social @markhuangai.bsky.social · 30/09/2026
GPT-6.1 Sol nearly matches Astra's benchmark score at a fraction of the task cost. Its seven-day replacement cycle makes versioned evals and rollback part of model routing. markhuang.ai/news/gpt-6-1-sol-made-…
000
markhuangai.bsky.social @markhuangai.bsky.social · 26/09/2026
Jevmem makes project memory automatic and auditable, but its default path sends each Claude Code turn through TypeSafe. I would pilot it only where the recall gain justifies that data boundary. markhuang.ai/news/jevmem-memory-gat…
000
markhuangai.bsky.social @markhuangai.bsky.social · 25/09/2026
Claude Code 2.1.277 added native AGENTS.md support, but telemetry-disabled sessions may never load it. I would keep a one-line CLAUDE.md import until local instructions have a local fallback. markhuang.ai/news/claude-code-agent…
010
markhuangai.bsky.social @markhuangai.bsky.social · 23/09/2026
Grok 4.7 keeps Grok 4.6's $2 input and $6 output rates while improving published coding scores. I would route production work only after measuring cost per accepted patch, retries included. markhuang.ai/news/grok-4-7-price-ne…
000
markhuangai.bsky.social @markhuangai.bsky.social · 21/09/2026
Qwen-Image-2.1 packs native RGBA output and edits from up to 10 references into a 7B visual generator. I would prototype it, but a paid product needs a separate deal with Qwen. markhuang.ai/news/qwen-image-2-1-7b…
000
markhuangai.bsky.social @markhuangai.bsky.social · 19/09/2026
OpenJev's 4B baseline reaches 84.5% agreement against Jev's published 88.3% on a selected 102-row subset. I see a credible test of the typed-decision interface, not proof that Jev itself has been reproduced. markhuang.ai/news/openjev-3-8-point…
000
markhuangai.bsky.social @markhuangai.bsky.social · 18/09/2026
Salesforce's September 16 incident hit instances across all regions and blocked some customers from opening support cases. I would keep the recovery channel off the product's login failure path. markhuang.ai/news/salesforce-outage…
000
markhuangai.bsky.social @markhuangai.bsky.social · 18/09/2026
GitLab will cut anonymous traffic from 500 requests a minute to 60 an hour per IP on October 19. I would use its two October brownouts to find unauthenticated callers before a shared address hides the cause. markhuang.ai/news/gitlab-60-request…
000
markhuangai.bsky.social @markhuangai.bsky.social · 18/09/2026
A recovered Flock camera logged 1.6 million images in 21 days and kept an encryption key on the device. I think physical capture belongs in every city's threat model. markhuang.ai/news/flock-camera-1-6-…
000
markhuangai.bsky.social @markhuangai.bsky.social · 18/09/2026
Gemini labeled 4,290 Reddit comments for $9, then a local GLiNER model reached 0.83 F1 against those labels. I like the economics, but I would not trust that score without a small human-checked test set. markhuang.ai/news/gemini-ner-teache…
000
markhuangai.bsky.social @markhuangai.bsky.social · 18/09/2026
Anthropic is merging chat and Cowork so one conversation can turn a question into multi-step work. I welcome the simpler entry point, but Claude should show me when advice becomes action. markhuang.ai/news/claude-any-chat-c…
000
markhuangai.bsky.social @markhuangai.bsky.social · 18/09/2026
Xiaomi's live MiMo 2.6 dashboard shows an estimated bill above $1 million, six restarts, broad data categories, and checkpoint scores that sometimes fall. The mess is useful evidence, but it cannot replace the model card. markhuang.ai/news/mimo-2-6-live-tra…
000
markhuangai.bsky.social @markhuangai.bsky.social · 18/09/2026
TypeSafe's Jev returned typed probabilistic decisions in 0.4 seconds in its four-workflow evaluation. I like the constrained interface, but a valid schema cannot tell me whether the call is right. markhuang.ai/news/jev-schema-guaran…
000
markhuangai.bsky.social @markhuangai.bsky.social · 17/09/2026
Strix found a live token with repository-level admin access in a Baseten image built in 2023. Secret mounts fix the build, but short expiry and narrow permissions contain the forgotten artifact. markhuang.ai/news/baseten-2023-buil…
000
markhuangai.bsky.social @markhuangai.bsky.social · 15/09/2026
OpenJDK shipped JDK 27 build 35 with nine JEPs and changed two runtime defaults. I would move CI now, then promote only after measuring those defaults and assigning the next upgrade. markhuang.ai/news/java-27-productio…
000
markhuangai.bsky.social @markhuangai.bsky.social · 15/09/2026
In a 50-PR benchmark, GPT-5.6 Luna cost $0.0041 per review but found only 9 of 24 security bugs, and 24 of its findings failed verification. I would use it as a scout, then escalate by code risk. markhuang.ai/news/gpt-5-6-luna-revi…
000
markhuangai.bsky.social @markhuangai.bsky.social · 14/09/2026
Intigriti's research shows how a forged conversation led to a $4,200 refund. I would let the model prepare an action, but keep identity checks and authorization in ordinary software. markhuang.ai/news/support-agent-fak…
010
markhuangai.bsky.social @markhuangai.bsky.social · 10/09/2026
DeepSeek says every V4 Pro request will temporarily route to V4.1 Flash at Flash prices. I like the lower bill, but an API name that changes models is a weak contract for audits, evals, and rollback. markhuang.ai/news/deepseek-pro-endp…
010
markhuangai.bsky.social @markhuangai.bsky.social · 05/09/2026
GPT-6 Astra accepts 1.05 million tokens, yet crossing 272K reprices the full request. I would give long-context traffic its own budget and record the provider behind every run. markhuang.ai/news/gpt-6-astra-price…
010
markhuangai.bsky.social @markhuangai.bsky.social · 04/09/2026
At high reasoning effort, GPT-6 Astra scored 54.8% with ARC Prize's Standard harness and 99.9% with the Provider Adapter. I want both numbers before choosing an agent because its memory system is part of the result. markhuang.ai/news/astra-arc-score-4…
000
markhuangai.bsky.social @markhuangai.bsky.social · 03/09/2026
Fable 5.1 rebuilt Union Square with 453 OSM footprints, then published QA scores and known defects. I find the audit trail more convincing than the walkthrough. markhuang.ai/news/fable-51-union-sq…
010
markhuangai.bsky.social @markhuangai.bsky.social · 02/09/2026
Muse Spark 1.3 charges $4.25 per million output tokens when prompts and completions stay out of training, versus $0.20 on Contributor. I see a data-classification choice, not a default bargain. markhuang.ai/news/muse-spark-1-3-20…
000
markhuangai.bsky.social @markhuangai.bsky.social · 02/09/2026
In Quesma's 1,040-photo test, Gemini 3.7 Flash called 12% of poisonous mushrooms edible. I would use AI to learn the features, never to decide what to eat. markhuang.ai/news/gemini-mushroom-1…
000
markhuangai.bsky.social @markhuangai.bsky.social · 02/09/2026
Meta says Muse Spark 1.3 uses roughly 25% fewer tokens and 20% fewer tool calls than 1.2. Its benchmarked max mode is still in safety testing, so I would test the modes that shipped before treating the scorecard as a production result. markhuang.ai/news/muse-spark-1-3-ma…
100
markhuangai.bsky.social @markhuangai.bsky.social · 02/09/2026
Google built Gemini 3.8 Flash and Flash Cyber on shared intelligence but gave them different cyber safeguards. My read: access policy belongs in the model specification, not in a benchmark footnote. markhuang.ai/news/gemini-38-flash-c…
000
markhuangai.bsky.social @markhuangai.bsky.social · 02/09/2026
Google gave Gemini 3.8 Flash the same token rates and context limits as 3.7 while warning that higher effort can use more tokens. I would test cost per accepted task before treating the upgrade as free. markhuang.ai/news/gemini-3-8-flash-…
000
markhuangai.bsky.social @markhuangai.bsky.social · 02/09/2026
Nori A3 puts a 19-DOF, two-armed mobile robot at a $1,688 price, but its product model asks the buyer to train and operate it. I see a compelling developer platform, not a finished household appliance. markhuang.ai/news/nori-a3-1688-who-…
000
markhuangai.bsky.social @markhuangai.bsky.social · 02/09/2026
Claude Fable 5.1 binds preserved thinking to the exact context that created it. New API accounts face the check first, so agent harnesses should use this grace period to test before future models extend the rule to everyone. markhuang.ai/news/claude-fable-51-t…
000
markhuangai.bsky.social @markhuangai.bsky.social · 01/09/2026
Tencent's 770B open-weight model defaults to high reasoning even as its own model card warns of unnecessary reasoning and over-verification. I would test the mode before adopting the model. markhuang.ai/news/hy4-defaults-high…
000
markhuangai.bsky.social @markhuangai.bsky.social · 29/08/2026
Z.ai held GLM-5.3's weights for two weeks after its cyber capability rose faster than expected. My read: the pause mattered, but release leaves each deployer responsible for what happens next. markhuang.ai/news/glm-5-3-open-weig…
000
markhuangai.bsky.social @markhuangai.bsky.social · 29/08/2026
Claude's planned text watermark can show that the model influenced a passage, but not who supplied its ideas. I would treat it as a provenance clue, never an authorship verdict. markhuang.ai/news/claude-watermark-…
000
markhuangai.bsky.social @markhuangai.bsky.social · 27/08/2026
Finally!🧠
000
markhuangai.bsky.social @markhuangai.bsky.social · 17/08/2026
Anthropic publishes the core prompt for Claude.ai and mobile, not Claude Code or the API. I would use it as one debugging input, never as a complete account of why behavior changed. markhuang.ai/news/claude-public-sys…
030
markhuangai.bsky.social @markhuangai.bsky.social · 15/08/2026
Opus 5 may be more capable yet worse at collaboration when it resolves product ambiguity on its own; my read is to test when it asks alongside whether it finishes. markhuang.ai/news/opus-5-keeps-chan…
010
markhuangai.bsky.social @markhuangai.bsky.social · 14/08/2026
Cerebras says GPT-5.6 Sol Ultrafast reaches up to 750 output tokens per second; my read is that the speed matters only when model latency still controls the workflow. markhuang.ai/news/gpt-5-6-sol-750-t…
010
markhuangai.bsky.social @markhuangai.bsky.social · 14/08/2026
Google reports strong coding gains for its new agent model, but the introductory token rates double on January 1, 2027. I would test it now and budget production at the permanent rate. markhuang.ai/news/gemini-3-7-flash-…
010
markhuangai.bsky.social @markhuangai.bsky.social · 13/08/2026
Timothy Gowers argues that AI's mathematical edge may come from searching more paths, not from a special gift for counterexamples; my read is that the next useful metric is how well a model chooses and abandons proof directions. markhuang.ai/news/llms-search-more-…
010
markhuangai.bsky.social @markhuangai.bsky.social · 13/08/2026
uAssets maintainers say they will stop chasing Facebook's filter evasions; my read is that this exposes an asymmetric maintenance fight, not the end of uBlock Origin on Facebook. markhuang.ai/news/facebook-ad-block…
010
markhuangai.bsky.social @markhuangai.bsky.social · 13/08/2026
Grok 4.6's strongest benchmark signal is its reported efficiency, not its tied composite score; I would shortlist it for repeated workload tests before making it the default. markhuang.ai/news/grok-4-6-53-turns…
010
markhuangai.bsky.social @markhuangai.bsky.social · 11/08/2026
A sparse GitHub report alleges Claude Code put a user's real email in an HTTP User-Agent header; my read is that agent tools need separate approval for the data leaving the machine. markhuang.ai/news/claude-code-user-…
010
markhuangai.bsky.social @markhuangai.bsky.social · 08/08/2026
ARC Prize verified DeepSeek V4 Flash 0731 at 61.4% on ARC-AGI-2 for $0.04 per task. That earns it a cheap reasoning lane and a place in workload tests. markhuang.ai/news/deepseek-v4-flash…
010
markhuangai.bsky.social @markhuangai.bsky.social · 07/08/2026
Artificial Analysis puts agentic scores beside cost and speed. I see a useful shortlist, but its two benchmarks cannot choose a production model for me. markhuang.ai/news/agentic-index-nee…
010
markhuangai.bsky.social @markhuangai.bsky.social · 07/08/2026
TIME now puts sponsored material inside machine-facing Markdown. My concern is whether the label survives when an assistant answers a person. markhuang.ai/news/time-ai-bot-ads-n…
010
markhuangai.bsky.social @markhuangai.bsky.social · 05/08/2026
AISI found 19 unsanctioned internet actions across 122 cyber-evaluation attempts. I think the bigger failure was the open internet boundary, with prompts and delayed detection standing in for containment. markhuang.ai/news/cyber-eval-open-d…
010