Sign in

camel-cdr.bsky.social

@camel-cdr.bsky.social
39 followers 53 following 117 posts

🐘 @camelcdr@tech.lgbt

PostsRepliesMedia
camel-cdr.bsky.social @camel-cdr.bsky.social · 23/10/2025
I only slightly disagree with using segmented load/store transpose. If you need to transpose from memory fine, but if you need register to register going though memory isn't the best. I'd use vslide1up/down or in the future vpaire/vpairo: github.com/ved-rivos/ri...
github.com
riscv-isa-manual/src/zvzip.adoc at zvzip · ved-rivos/riscv-isa-manual
RISC-V Instruction Set Manual. Contribute to ved-rivos/riscv-isa-manual development by creating an account on GitHub.
010
camel-cdr.bsky.social @camel-cdr.bsky.social · 23/10/2025
"How NOT To Program an Out-of-order Vector Processor" slides are public. static.sched.com/hosted_files...
111
camel-cdr.bsky.social @camel-cdr.bsky.social · 12/10/2025
Fuzzing tip: use VLA instead of fixed-size buffers or malloc 1. with fixed-size buffers asan won't catch everything. 2. VLAs are faster than malloc, in my case I get 15% faster fuzzing. If VLAs aren't portable enough, just check __STDC_NO_VLA__ and select between the other options.
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 25/09/2025
*correction: 0.5/0.5/2/4 for vector-scalar/immediate compares (0.5/2/4/8 for vector-vector)
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 25/09/2025
For the scalar instructions: * 6-issue: add/sub/lui/xor/sll/shNadd/zext/clz/cpop/min/rotl/rev8/bext/... * 3-issue: load/store * 2-issue: fadd/fmul/fmacc/fmin/fcvt * 1-issue: mul/mulh/feq/flt * pipelined: fsqrt/fdiv: ~8.5, div/rem: 12-16
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 25/09/2025
My takeaway so far is to not be scared to use the segmented load/stores, and LMUL>1 permutes are good, but you probably want to avoid LMUL=8 ones when possible. I'll continue manually unrolling none-lane-crossing permutes. For LMUL>1 comparisons, it's better to use .vx/vi over .vv when possible.
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 25/09/2025
The vslide1up/vslide1down do scale perfectly, though, with 0.5/1/2/4. It's not in the benchmark, but I hope vslideup/vslidedown with immediate "1" also do. We'll have to wait for the other microbenchmarks to get a more complete picture.
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 25/09/2025
* Ovlt behavior isn't supported, but I don't really care much about it The only bigger negative thing I've seen so far is that the vslideup/vslidedown instructions don't scale linearly or close to linearly with LMUL, even for a small immediate shift amount like "3".
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 25/09/2025
* dual-issue vrgather, with good scaling: 0.5/1/8/30 * dual-issue vcompress, with OK scaling: 0.5/3/6/17 (I still think this could get close to linear) * Fault-only-first loads seem to have no overhead * Segmented load/stores look quite fast, even the more exotic ones like seg7
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 25/09/2025
* Most instructions have an inverse throughput of 0.5/1/2/4 for LMUL=1/2/4/8, even vslide1up/down, 64-bit vmulh, viota, vpopc and integer reductions * 0.5/0.5/1/2 for vector-scalar/immediate compares and 0.5/1/2/- for narrowing instructions (see "Microarchitecture speculations" section)
200
camel-cdr.bsky.social @camel-cdr.bsky.social · 25/09/2025
Tenstorrent decided to publish the first benchmark data for Ascalon's RVV implementation using the instruction throughput benchmark of my rvv-bench benchmark suite. <3 camel-cdr.github.io/rvv-bench-re... Overall, the results look really good so far:
133
camel-cdr.bsky.social @camel-cdr.bsky.social · 24/08/2025
Third Way is an unfortunate name: en.wikipedia.org/wiki/Third_W...
000
Reposted by @camel-cdr.bsky.social
Claire Xen 🏳️‍⚧️ 🧙🏻‍♀️ 💖💛💙 @clairexen.bsky.social · 26/07/2025
So if you are currently involved with ISA-level decisions about inclusion of any pext/pdep-like instructions: Please consider including SAG/inverse-SAG with bit-reversal of the goats. No matter which of the two implementation methods you are using: All you need to do is not mask the goat bits.
043
camel-cdr.bsky.social @camel-cdr.bsky.social · 24/07/2025
It looks like the patent expires at the end of 2028. The earliest I could see a RVI extension ratified at this point is 2027, so it's definitely worth evaluating. Also, the new diagrams are cool.
110
camel-cdr.bsky.social @camel-cdr.bsky.social · 20/07/2025
I've only watched the last hour this far, but I quite liked your take on null-terminated strings. C really has to be understood with its history in mind.
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 11/07/2025
www.youtube.com/watch?v=OPgj...
youtube.com
Ventana’s Second Gen RISC V Processor for Data Center and Other High Performance | Greg Favor
YouTube video by Ventana Micro
010
camel-cdr.bsky.social @camel-cdr.bsky.social · 11/07/2025
Their V2 slides say, that they have a macro-op cache equivalent in size to a regular 32 KiB icache. It can store variable length entries of up to 48 macro ops, which can be fuses from non-sequential instruction runs by collapsing taken branches.
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 11/07/2025
TIL about Trace Cache: www.realworldtech.com/forum/?threa... (thread on Apples Trace Cache) Ventanas Veyron V2/V3 seem to also use something like a trace cache.
realworldtech.com
RWT Forums - Real World Tech
content overridden
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 28/06/2025
Ohh, the talk recordings are on YouTube: www.youtube.com/watch?v=1lwz...
youtube.com
CBP2025 - Opening Remarks - Rami Sheikh
YouTube video by Rami Sheikh
010
camel-cdr.bsky.social @camel-cdr.bsky.social · 28/06/2025
The sixth Championship of Branch Prediction (CBP2025) happened a week ago: ericrotenberg.wordpress.ncsu.edu/cbp2025-work...
152
Reposted by @camel-cdr.bsky.social
Claire Xen 🏳️‍⚧️ 🧙🏻‍♀️ 💖💛💙 @clairexen.bsky.social · 20/06/2025
I wrote a reference implementation for a SAG without bit reflection: github.com/clairexen/ed..., and I wrote a parametric SAG core for any bit width: github.com/clairexen/ed...
github.com
edu-sag/param.v at main · clairexen/edu-sag
Educational 8-Bit Sheep-And-Goats (SAG) Verilog Reference IP - clairexen/edu-sag
012
camel-cdr.bsky.social @camel-cdr.bsky.social · 16/06/2025
>>> lut=np.array([ord('a'),0,ord('e'),0,ord('i'),0,0,ord('o'),0,0,ord('u'),0,0,0,0,0], dtype=np.uint8) >>> inp=np.frombuffer(b"test128aeiou72761xjs",dtype=np.uint8) >>> lut[(inp&31)>>1]==inp
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 12/06/2025
4x 16-bit: 120 u^2 63% utilized, 5GHz met (49 slack) 2x 32-bit: 120 u^2 65% utilized, 5GHz met (52 slack) 1x 64-bit: 153 u^2 64% utilized, 5GHz met (14 slack) So subsetting on SEW really doesn't make much sense compared to a .vx subset.
110
camel-cdr.bsky.social @camel-cdr.bsky.social · 12/06/2025
I got OpenROAD working and tested the bfly part of your implementation (so without decode) in a SIMD setup. asap7, targeting 5GHz, 75% placement density and 50% utilization:
110
camel-cdr.bsky.social @camel-cdr.bsky.social · 06/06/2025
SiFive X280 RVV benchmarks: camel-cdr.github.io/rvv-bench-re... Civil was so nice run my RVV benchmark on the SiFive X280 cores on the Tenstorrent Blackhole.
camel-cdr.github.io
RVV benchmark SiFive X280
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 06/06/2025
I just had this problem on RISC-V where I didn't clobber the vector registers and some autovectorized surrounding code broke on a newer kenel version.
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 06/06/2025
TIL you can't do forward compatible syscalls with inline assembly because the kernel can decide to clobber architectural state that was added after you wrote the code. If you use svc with inline assembly, you have to explicitly clobber SVE registers. Good luck doing this back in 2015 when you wrote
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 06/06/2025
I suppose, the instruction encoding space has to be considered as well.
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 06/06/2025
Ah, I understand my mistake now. My mental model had the element order between the stages as fixed, which is why I didn't see the equivalence of the graphs.
010
camel-cdr.bsky.social @camel-cdr.bsky.social · 06/06/2025
Guess I'll step up: github.com/camel-cdr/bf... And, yes, I wasn't the first person to write an optimizing brainfuck interpreter in the c preprocessing, that honor goes to kotha.
github.com
010
camel-cdr.bsky.social @camel-cdr.bsky.social · 06/06/2025
Also, I think there should be a compress_right_flip and compress_right instruction, because most cases would want the mask and it's basically free to add in hardware.
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 06/06/2025
I (not a hardware person) would expect a regular ibfly to have 2+4+8 parallel wires, while this approach has 10+10+10. (both +4 for every layer, if you include the control signals) For 8-bit this is quite tame, but for 64-bit this may be a significant factor.
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 06/06/2025
Ah, I think I understand this better now. The unshuffle approach is quite neat, but isn't it a lot worse in terms of wire crossings/area/delay for a single cycle implementation?
200
camel-cdr.bsky.social @camel-cdr.bsky.social · 05/06/2025
I thought a butterfly network can't do SAG in one pass: programming.sirrida.de/bit_perm.htm... > […] sheep-and-goats operation generally cannot be performed on a butterfly network; […] a variant thereof can which gathers the […] bits on one end in order, but also the remaining ones mirrored […]
programming.sirrida.de
Bit permutations
An essay about bit permutations in software
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 04/06/2025
I think something like vrgather128ei4.vx would be useful. It would do a 16-byte shuffle in all 128-bit lanes controlled by the nibbles in the GPR. This could use the same shuffle hardware as vrgather, but scales better with LMUL/large vlen and saves a vector register allocation.
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 04/06/2025
So like xperm4? That would require new, presumably expensive, bit shuffle hardware. With vbmatflip+vrgather+vbmatflip you could shuffle bits in bytes and using vrgather+vbmatflip+vrgather+vbmatflip+vrgather gives you bits in 16b/32b/64b/... permute (the limit depends on VLEN).
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/06/2025
On that topic. I think a 16-bit bit permute would be extremly cheap to implement. You just chain the bfly and ibfly networks and need 2x 4x8-bit = 64 control bits, which fit in a GPR. Something like vb16perm.vx vrd, vs, rs. I don't see any chances getting that through RVI though, to impl specific
200
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/06/2025
Looking forward to it.
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/06/2025
SAG would be nice, if it's cheap with the pext/pdep hardware, but I'm not sure I follow the "not mask the unselected bits". Wouldn't you run the bfly and ibfly in parallel and combine the results, which would require a new shifter. Or implement compress-flip, which needs a partial bit reverse.
200
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/06/2025
Sidenote: My pseudocode for the LEB128 decoder using RVV pext/pdep instructions isn't completely correct. I'll revisit it properly, with spike/qemu implementation, once I finish my project.
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/06/2025
Lastly, do you know any other bitmanip instructions that were considered and may be useful for RVV? This also goes to everybody else reading; I'd love to hear interesting SIMD instruction suggestions.
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/06/2025
This would mean implementations of that subset would only have to decode one mask to control signals; it might even be possible to do the prefix sum on the viota.m hardware. Is implementing bmatflip via chaining the bfly and ibfly network uses in the pdep/pext implementation reasonable?
110
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/06/2025
The original HILEWITZ thesis mentioned that only the decoder needed to be pipelined and that the decoder took most of the area. Is this the same in your implementation? I think it would be worth having a subset of the full extension with only a .vx instruction variant (mask from GPR for elements).
210
camel-cdr.bsky.social @camel-cdr.bsky.social · 03/06/2025
@clairexen.bsky.social Hi Claire, we are trying to propose some of the dropped bitmanip instructions for RVV: lists.riscv.org/g/sig-vector... Since you were deeply involved in the development of the bitmanip spec, I was wondering if you could answer some questions about your bextdep implementation.
lists.riscv.org
[Proposal] Bit Compress & Bit Decompress Instructions for RVV
For Spark/Flink workloads in data centers, reading large-scale Parquet files is often a performance bottleneck. Therefore, adding support for these instructions can effectively fill this gap, ensuring RISC-V's competitiveness with other ISAs.
210
camel-cdr.bsky.social @camel-cdr.bsky.social · 26/05/2025
Ok, there don't seem to be any bugs related to this in dav1d.
000
camel-cdr.bsky.social @camel-cdr.bsky.social · 26/05/2025
Edit: I thought I found a dav1d bug (vwadd.wx v0, v0, v8), but I didn't norice the .wx, so it wasn't a bug. I'll have to check the rest of the code later.
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 26/05/2025
looks like gcc generates wrong code, and clang is to conservative with overlaps and generates redundant moves: godbolt.org/z/1czr8oGab just created a bug report: gcc.gnu.org/bugzilla/sho... I'll have to check all RVV assembly I've written.
100
camel-cdr.bsky.social @camel-cdr.bsky.social · 26/05/2025
oh no > When source and destination registers overlap and have different EEW, the instruction is mask- and tail-agnostic, regardless of the setting of the vta and vma bits in vtype.
When source and destination registers overlap and have different EEW, the instruction is mask- and tail-agnostic, regardless of the setting of the vta and vma bits in vtype.
120
camel-cdr.bsky.social @camel-cdr.bsky.social · 13/05/2025
"Efficient Implementation of RISC-V Vector Permutation Instructions" -- arxiv.org/abs/2505.07112 "Efficient Architecture for RISC-V Vector Memory Access" -- arxiv.org/abs/2504.08334 I love how these two were released so close to each other.
031
camel-cdr.bsky.social @camel-cdr.bsky.social · 13/05/2025
010