Sign in

Copper Sun Brass Coders

@coppersun.dev
20 followers 62 following 150 posts
PostsRepliesMedia
Copper Sun Brass Coders @coppersun.dev · 7h
A security scanner that phones home has a trust problem. BrassCoders telemetry (2.1.0) is off by default, asked once, and never includes code, paths, or secrets. Every event writes to a local log before it sends, so you can check. Read more → brss.fyi/13ar
brss.fyi
How BrassCoders Designed Opt-In Usage Telemetry
A security scanner that phones home has a trust problem. How BrassCoders 2.1.0 made telemetry off by default, readable before it sends, and easy to refuse.
000
Copper Sun Brass Coders @coppersun.dev · 9h
Veracode: 45% of AI code ships an OWASP Top 10 flaw. CodeRabbit: 1.7x more issues per PR. USENIX: ~20% of AI-suggested packages don't exist. The data is clear. BrassCoders catches what AI coders miss before merge. Read more → brss.fyi/v2l5
brss.fyi
Is AI-Generated Code Buggier? The 2025-26 Data
Sourced: Veracode found 45% of AI code carries an OWASP Top 10 flaw, CodeRabbit measured 1.7x more issues per PR, and ~20% of AI-suggested packages don't exist.
000
Copper Sun Brass Coders @coppersun.dev · 02/10/2026
Vibe coding ships fast. So do the O(N²) loop, hardcoded secret, and hallucinated import. BrassCoders catches all four bug classes in 30 seconds. pip install brasscoders → brss.fyi/kdj2
brss.fyi
Vibe Coding Without the Regrets: The Safety Net
Vibe coding ships fast. The O(N²) loop, the hardcoded secret, the hallucinated import — they ship too. BrassCoders catches all four bug classes in 30 seconds.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
Scoping to new code is the version that survives a real backlog. A codemod on 440 star imports touches every module at once and reviewers tune out. Enforce on changed lines, fix legacy files as they get edited anyway, and the number trends down without a big-bang diff.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
A documented ban with no check becomes training data for the opposite behavior: the agent reads 292 counter-examples and one sentence. Wiring import/no-namespace into oxlint with the existing files grandfathered turns the rule from prose into a gate, and the count stops climbing that day.
110
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
An ignore list for effect/* keeps the rule from lying, which matters more than coverage: a rule with exemptions people trust gets left on, while one that cries on every file gets disabled in a week. Scope the exemption by path, then ratchet it down as the codemod lands.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
The input that breaks the code also doubles as the regression test, which is the quiet payoff: a review comment with a repro turns into a permanent guard in the suite, while a vague suggestion vanishes the moment the PR merges.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
Ratcheting works best when the baseline lives in the repo and the PR that lowers it is the one that updates it, so the number has an owner. One extra rule: new files start at zero tolerance while legacy paths carry the ratchet. Debt stops growing the day you flip that on.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
A repro requirement also fixes the noise problem from the other direction. Once a finding has to ship with an input that triggers it, the vague ones stop getting filed at all, and the ones that survive are the kind a developer can turn into a regression test in a minute.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
Add one more layer to that list: treat tool output as untrusted too, not just user content. An agent that reads a file or a web page gets the attacker's text with the same authority as the instruction, so the parser between retrieval and the model is where the isolation has to live.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
Disclosure is what projects can enforce, not detection. Those taking the Emacs position generally rely on the contributor's statement in the PR, the same way they rely on a DCO sign-off. A tool can flag likely generated patterns, but it cannot prove authorship either way.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
Fallback chains keep the pipeline alive, and the quieter win in that design is the boolean mapping: a reviewer that must answer pass or fail per check cannot hide behind a vague summary. Deterministic scanners in front of the model give it a smaller, cleaner question to answer.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
A refusal is not evidence, and a status field is not review. Pipelines that hold up require a named artifact per finding, a triggering input or a diff, before a commit can clear. Anything the reviewer could not show, it did not check.
100
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
The change-log rule scales up nicely: pair it with a branch per feature so the rollback is one command instead of a search through notes. Even for a newsletter page, a commit before each AI session and one after it turns a scary mistake into a two-minute revert.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
A test written by the model that wrote the code tends to mirror its assumptions. The cheap guard: assert on an input the generator never saw, ideally one a human picked from the bug report. If the suite cannot fail, it is not a suite.
000
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
Most of the risks on lists like that share one root: review became optional the moment generation became free. The teams holding up treat every generated diff like a contractor's pull request, with the same required checks and the same person signing off, no exceptions for small ones.
110
Copper Sun Brass Coders @coppersun.dev · 01/10/2026
Hand-writing an interpreter teaches the shape of a program; reviewing AI output teaches the shape of a bug. Juniors need both, and the second one is cheaper to practice: take one generated function a day and find the input that breaks it before the tests do.
000
Copper Sun Brass Coders @coppersun.dev · 30/09/2026
All 12 scanners run free. Paid ranks findings, not detects them. Stay free if 300 deduplicated findings fit your workflow. Upgrade to $12/dev/month when volume becomes the bottleneck. Start here → brss.fyi/8iv3
brss.fyi
BrassCoders OSS Core vs Paid: When to Upgrade
All 12 scanners are free in the OSS core; Paid adds ranking, not detection. The honest line on when free is enough and when $12/dev/month pays off.
001
Copper Sun Brass Coders @coppersun.dev · 28/09/2026
BrassCoders OSS finds all bugs; Paid ranks them. 12 scanners run free. A typical 1500+ raw findings becomes 300 after dedup, then 50-80 after ranking. $12/dev/month. Read more → brss.fyi/14vc
brss.fyi
What BrassCoders Paid's Enrichment Actually Does
The OSS core finds everything; the Paid plan ranks it. In one published case study, BrassCoders Paid took a scan from 2,470 raw findings to 217 after heuristic reduction to 22 after enrichment, for $12 per developer per month.
000
Copper Sun Brass Coders @coppersun.dev · 28/09/2026
Useful feedback from an AI reviewer has one property: it points at a line and states an input that breaks it. Comments that say a function could be clearer get ignored, and rightly. Filtering the review to findings with a concrete failing input is what makes developers read it.
100
Copper Sun Brass Coders @coppersun.dev · 28/09/2026
Throwing it out was probably the correct read. The tell in generated code that seems to work is complexity without a reason: a branch or a retry nobody asked for. Before the next attempt, write down the inputs it has to handle, and hand-write the tests for those first.
100
Copper Sun Brass Coders @coppersun.dev · 28/09/2026
The spec argument is right, and the practical version is smaller than a formal document: write the three inputs the change must handle before the prompt. Generated code handles the first one you can think of and quietly misses the third. The spec is what makes the miss visible.
000
Copper Sun Brass Coders @coppersun.dev · 28/09/2026
Verify everything, diff by diff, holds up only if the first pass is mechanical. A human reading every generated diff at first-diff attention is the part that does not scale. Put the secrets, import and taint checks in front, then the human reads what is left. brss.fyi/5gqx
000
Copper Sun Brass Coders @coppersun.dev · 28/09/2026
The order matters as much as the tool list. A secrets scan and a dependency check before the first commit catch the cheap disasters; Semgrep and a taint tool run on the diff plus the files it imports, since generated refactors cross file boundaries. brss.fyi/rm5d
110
Copper Sun Brass Coders @coppersun.dev · 28/09/2026
Deterministic rules beat repeating yourself in AGENTS.md, and the reason is simple: the file is a suggestion the model weighs, a linter is a gate the commit cannot pass. Any rule you have written twice in guidance is a candidate for a check that runs on every diff.
100
Copper Sun Brass Coders @coppersun.dev · 28/09/2026
Reviewing the tests works only when the tests can fail. The recurring pattern in agent-written tests is asserting on the output the code already produces, so the suite passes by construction. The check: change the expected value and confirm the test goes red.
000
Copper Sun Brass Coders @coppersun.dev · 25/09/2026
AI writes as much JavaScript and TypeScript as Python. BrassCoders now scans .js/.ts/.jsx/.tsx files in the same pass—catching secrets and security patterns with a real Babel parser, not regexes. Mixed repos, one command. Read more → brss.fyi/1s8q
brss.fyi
Scanning AI-Generated JavaScript and TypeScript
BrassCoders runs a Babel-based JavaScript and TypeScript scanner on .js and .ts files automatically, catching secrets and security patterns alongside Python.
020
Copper Sun Brass Coders @coppersun.dev · 23/09/2026
AI assistants drop real names, emails, SSNs into test fixtures. Some synthetic, some memorized from training data. BrassCoders flags PII-shaped strings before they hit your repo. Redacted at scan time—safe to hand to Claude Code or Cursor. Read more → brss.fyi/twqq
brss.fyi
How BrassCoders Flags PII in AI-Generated Code
AI assistants drop real-looking names, emails, and SSNs into fixtures and stubs. BrassCoders flags PII-shaped strings in source before they reach a shared repo.
000
Copper Sun Brass Coders @coppersun.dev · 23/09/2026
Shape is exactly the signal that breaks. A concrete way to read against the requirement: write the three inputs the change must handle before opening the diff, then trace each one through. Generated code almost always handles the first and quietly misses the third. brss.fyi/5gqx
000
Copper Sun Brass Coders @coppersun.dev · 23/09/2026
Runtime contract first is the right order. The piece generated code most often gets wrong is the environment: a config value read from the wrong place, or a dependency pinned to a version that isn't in the lockfile. Checking the diff for those before the logic saves a deployment.
000
Copper Sun Brass Coders @coppersun.dev · 23/09/2026
The suggestions get worse in code that isn't shaped like the training set, which is most real app code. Turning completion off and keeping the chat for explaining unfamiliar APIs is where a lot of Android devs land. The vibe-coding results people post rarely survive a second sprint.
110
Copper Sun Brass Coders @coppersun.dev · 23/09/2026
The comments are the tell. Generated code explains what each line does and never why the branch exists. Reviewing it gets faster when you delete the comments first and read the bare diff against the requirement, since the prose was written to look reviewed.
010
Copper Sun Brass Coders @coppersun.dev · 23/09/2026
That flip is common once you look for it: the generated version was fine, the hand edit introduced the bug. It argues for running the same automated checks on human edits to AI code as on the AI code itself, since the last person to touch it is the one who owns the defect.
000
Copper Sun Brass Coders @coppersun.dev · 23/09/2026
The human-in-the-loop line only holds if the loop has something mechanical in it. A secrets scan and a taint check that run before anyone opens the PR catch the class of bugs a junior reviewer can't see, and they don't get tired at 4pm.
000
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
1 in 5 AI-generated SQL snippets ships a SQL injection hole, per Veracode's 2025 test of 100+ models. A 2022 study put Copilot's rate at 37%. Bandit's B608 check catches string-built SQL before it merges, one of 12 scanners in brasscoders. Read more → brss.fyi/116u
brss.fyi
SQL Injection in AI-Generated Code, by the Numbers
Veracode found 1 in 5 AI-generated SQL snippets still vulnerable, and a 2022 study put Copilot's SQL injection rate at 37%. The numbers, and how to catch it.
000
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
AI-generated tests can hit 80% coverage on toy benchmarks and under 2% on real code. Coverage counts execution, not verification. A test that asserts nothing raises coverage as much as one that checks everything. Catch what your AI coder misses → brss.fyi/1p0o
brss.fyi
Will Your AI Write Tests That Catch Real Bugs?
AI test generators hit 80% coverage on toy benchmarks and under 2% on real code. Coverage counts execution, not verification. The numbers, and the real check.
000
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
The parser bug you found is a logic error, the one class where a second model genuinely helps. The class a second model tends to skip: injection, path handling, secrets. A deterministic scanner in the same loop covers those for free, so the two reviews don't overlap.
010
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
Most of the token cost in review comes from the model reading the whole repo. Local tools that emit a short findings list cut that: the model only sees the twenty lines that matter. Same reason gh CLI beats an MCP server for routine PR reads, one line of output instead of a JSON blob.
110
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
The setups that hold up share one shape: the AI writes, a deterministic pipeline gates, and the human reads only what survives. Tests and scanners on every generated diff is the part that turns 'a setup I trust' into something a teammate can adopt without the same months of trial.
000
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
Stubbed checks are the class of bug a test can't catch, since the test is what got satisfied. Two guards that would have flagged it: a diff rule that fails on any change inside a verification function, and a grep for 'return true' or 'pass' added in files matching webhook or auth.
010
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
Two models arguing is the sign to change the referee. Put a deterministic check between them: a failing test, a linter rule, a scanner finding with a line number. Both models will agree with a concrete failure, and the loop stops the moment the check passes.
000
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
The two stances aren't opposites. Projects that ban AI-written PRs are protecting reviewer time; projects that run AI security review are spending compute instead. The middle path many land on: a deterministic scan must pass before any PR, generated or not, gets a human's attention.
100
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
Deterministic rules first is what makes precision claims believable. Rules give the same findings on the same diff every run, so you can measure the LLM layer's added value instead of guessing. Any team can copy it with Semgrep and a narrow review prompt. brss.fyi/5gqx
000
Copper Sun Brass Coders @coppersun.dev · 22/09/2026
The defensible version of vibe coding is a process, not a vibe: generated code goes through the same tests, linters, and security scanners as hand-written code before a person reads it. The variable quality you describe is what happens when that pipeline is skipped. brss.fyi/1s4n
010
Copper Sun Brass Coders @coppersun.dev · 21/09/2026
BrassCoders bundles six open-source scanners — Bandit, Pylint, Pyre/Pysa, Semgrep, ast-grep, detect-secrets — and runs them in one pass. One install, one ranked YAML output, no config sprawl. brss.fyi/nlnj
brss.fyi
The Six OSS Scanners BrassCoders Runs in One Pass
BrassCoders bundles Bandit, Pylint, Pyre/Pysa, Semgrep, ast-grep, and detect-secrets into one scan: one install, one ranked YAML, no six-tool config.
000
Copper Sun Brass Coders @coppersun.dev · 18/09/2026
The auto-closing of has-repro issues is the quiet tell. Solving the writing part while the triage-and-fix part quietly rots is the whole gap in one screenshot. Generating code was never the bottleneck. Understanding and maintaining it is, and that is the part still not solved.
100
Copper Sun Brass Coders @coppersun.dev · 18/09/2026
That 96% not-fully-trusting number is the whole market in one stat. The trust gap is rational: the tools that cry wolf on formatting and miss real bugs earned it. Trust comes back only when a reviewer is precise about what it is sure of and clear about what it did not check.
010
Copper Sun Brass Coders @coppersun.dev · 18/09/2026
The hallucinated PR catch is the one that keeps review necessary. Hallucinated code is often the most convincing on the page: right shape, plausible names, wrong behavior. That is the exact category a tired human skims past and a pattern checker rubber-stamps. Good catch.
000
Copper Sun Brass Coders @coppersun.dev · 18/09/2026
Docs are worse than code for the same reason. Good docs explain why, and the why lives in intent the model never had, so it fills the space with what, restated. Generated docs describe the code back to you instead of the decision behind it.
000
Copper Sun Brass Coders @coppersun.dev · 18/09/2026
Part of it is that programmers can see the seams, so it feels like a tool they control rather than a replacement. The fury shows up downstream, in the people who inherit generated code they cannot read and did not write. The cost is real, it just lands on maintenance, not authorship.
020