Sign in

Daniel Estévez

@destevez.net
692 followers 67 following 367 posts

Everything space & RF. Amateur radio operator (EA4GPZ / M0HXM). PhD in Mathematics from Univ. Autónoma de Madrid. he/him

PostsRepliesMedia
Daniel Estévez @destevez.net · 11/09/2026
In this recording, Roman is transmitting 19.5 kbaud PCM/PSK/PM with a subcarrier of 1.7 MHz, using CCSDS concatenated coding with 8 interleaved RS codewords, and AOS frames. I have managed to interpret the Space Packet timestamps and find some ASCII strings. Read more: destevez.net/2026/09/deco...
destevez.net
Decoding the Roman Space Telescope – Daniel Estévez
040
Daniel Estévez @destevez.net · 11/09/2026
New blog post: Decoding the Roman Space Telescope. On August 31 I observed the Roman telescope S-band telemetry with the Allen Telescope Array some time after launch. This post is a summary of the observation, including the telemetry decode and a preliminary analysis.
A waterfall of the signal of the spacecraft.
1124
Daniel Estévez @destevez.net · 06/09/2026
Read more: destevez.net/2026/09/on-m...
destevez.net
On matched filtering in the presence of carrier frequency offset – Daniel Estévez
020
Daniel Estévez @destevez.net · 06/09/2026
I close with some remarks about usage in practice. When the frequency grid spacing is narrower than the inverse of the number of samples, the conventional matched filter performs quite well and the DPSS-windowing improvement is small. For larger spacings it provides a minor gain.
A plot that compares the average SNR loss with respect ideal filtering of both the classical matched filter and DPSS-windowed matched filter, in terms of the frequency grid spacing.
110
Daniel Estévez @destevez.net · 06/09/2026
This is related to the spectral concentration problem, and the solution for constant-envelope waveforms is given by the DPSS window.
A plot that shows DPSS windows for different frequency grid spacings Delta. For small Delta the window is almost flat (like the rectangular window). As Delta increases, the window becomes more tapered.
100
Daniel Estévez @destevez.net · 06/09/2026
Concretely, I explain how to maximize the average SNR when a bank of filters spaced on a regular grid of frequencies is used and the carrier frequency offset is distributed uniformly.
100
Daniel Estévez @destevez.net · 06/09/2026
New blog post: On matched filtering in the presence of carrier frequency offset. In this short maths note I explain how to tweak a matched filter to improve the SNR when there is carrier frequency offset.
120
Daniel Estévez @destevez.net · 27/08/2026
Finally I go into the technical details of how to implement all this with the mlir-aie framework, and the challenges to program the DMA engines in the way we want. I close with some benchmarking results. Read more: destevez.net/2026/08/x-en...
destevez.net
X-engine correlator on a Ryzen NPU – Daniel Estévez
000
Daniel Estévez @destevez.net · 27/08/2026
Next I consider the data flow through the whole NPU array. This leads to a design of the Ethernet packets that the X-engine ingests. There is a tradeoff between having the data in the order that the X-engine prefers and having it in an order that is easy to generate in the PFBs.
100
Daniel Estévez @destevez.net · 27/08/2026
Then I implement and benchmark the C++ kernel that does the matrix multiplication (of some submatrices), and I discuss some factors that can cause performance loss, such as memory bank conflicts. Then I design how to use the compute tile memory and DMA data flows.
100
Daniel Estévez @destevez.net · 27/08/2026
In the blog post I go over all the details of the bottom-up design and implementation and all the challenges I found. I explain how to do some back-of-the-envelope calculations to estimate performance before we begin coding.
100
Daniel Estévez @destevez.net · 27/08/2026
The right platform to deploy this kind of X-engine would be a Versal AI Edge Gen2 FPGA, which has an AIE-ML v2 engine that is very similar to this Ryzen NPU. However, ingesting such a high data rate on these Versal parts is also challenging (the large parts have 3x 100GbE MACs).
100
Daniel Estévez @destevez.net · 27/08/2026
Testing this design on a Ryzen laptop is just for prototyping and learning. The X-engine ingests 65.5 GB/s and there is no way that I can input that amount of data from the outside world into a laptop (this Ryzen SoC has 16x PCIe 4.0, which would give 31.5 GB/s).
100
Daniel Estévez @destevez.net · 27/08/2026
My implementation handles 256 streams (which can be 128 dual polarization antennas) at 64 Msps IQ (aggregated rate over a number of PFB channels) in real time. The NPU uses 1W of power, and the whole Ryzen SoC uses 25W. It is amazing how much compute we get for that power.
110
Daniel Estévez @destevez.net · 27/08/2026
NPUs are essentially arrays of vector processors intended to perform large matrix multiplications very fast, and an X-engine does just that: a large matrix product. However, getting the X-engine running right on the NPU is trickier than it sounds, due to many low-level factors.
100
Daniel Estévez @destevez.net · 27/08/2026
New blog post: X-engine correlator on a Ryzen NPU. A few months ago I posted about the basics of NPUs in Ryzen AI chips. I wanted to tackle a real-world problem to get more experience, so I've implemented an X-engine correlator like the ones used in radio astronomy.
132
Daniel Estévez @destevez.net · 20/08/2026
As we're seeing more and more issues and PRs that use LLMs in some way, the gr-satellites project is considering adopting an AI policy to have clearer guidelines of what to expect and how to proceed. If you have strong thoughts, join the discussion in github.com/daniestevez/...
github.com
Request: Add an AI policy, maybe · Issue #849 · daniestevez/gr-satellites
Per the discussion in #846, it'd be helpful to have an AI usage policy for contributions in this repository.
130
Daniel Estévez @destevez.net · 07/08/2026
New blog post: Tianwen-1 safe mode telemetry. Peter from @amsat-dl.org noticed that a few days ago Tianwen-1 was transmitting in safe mode. Here I analyze the telemetry signal, which is only 32 baud (16 bits per second). destevez.net/2026/08/tian...
A screenshot of a GNU Radio GUI that shows the spectrum of the signal and a time domain plot of the demodulated symbols.
0101
Daniel Estévez @destevez.net · 30/07/2026
The Madrid "Sierra oeste" fire can be seen as a large purple region in the centre right of the image. On the upper left, the scar left by last year's Orense's fires can still be seen as lighter purple blotches. Smaller fires can be seen peppered across Spain and Portugal.
000
Daniel Estévez @destevez.net · 30/07/2026
Burnt areas appear in dark red or purple, due to the dryness causing lower water absorption in shortwave infrared, and higher absorption in near infrared caused by soot. Healthy vegetation appears in green due to the high reflectance of leaf cellular structures in near infrared.
100
Daniel Estévez @destevez.net · 30/07/2026
This Sentinel-2 infrared image from yesterday shows the extent of the Madrid wildfires, which are now stable and do not show any smoke in satellite imagery. The false colour map shows two shortwave infrared bands (B12 and B11) in red and blue, and near infrared (B8) in green.
A false colour satellite image that shows the western half of Spain and all of Portugal. The tweets describe how the false colour map works and what fires can be seen. Besides that, it is noticeable the higher amount of vegetation in the north and northwest part of the Iberian peninsula, and also in a few small regions in the interior.
111
Daniel Estévez @destevez.net · 17/07/2026
Here is the full job description: job-boards.greenhouse.io/muonspace/jo...
job-boards.greenhouse.io
Senior FPGA RF Signal Processing Engineer
Mountain View, CA or Remote
000
Daniel Estévez @destevez.net · 17/07/2026
Muon Space is looking for an FPGA engineer for RF signal processing. Join a team of FPGA engineers and RF DSP engineers including myself to build state of the art RF instruments. Ability to access ITAR/EAR information required. DM me if you are interested or have questions.
122
Daniel Estévez @destevez.net · 09/07/2026
In my solution I've gone fancy and implemented a Kalman filter (here I won't say exactly why to avoid spoilers). The CTF and solutions are available in Github, so anyone can play at their own pace. destevez.net/2026/07/oris...
destevez.net
ORI’s lunar descent CTF – Daniel Estévez
010
Daniel Estévez @destevez.net · 09/07/2026
New blog post: @openresearch.institute.web.brid.gy's Lunar descent CTF. This is my solution to a CTF that simulates the Ka-band radio altimeter of Chandrayaan-3. The goal is to understand and fix a bug in the measurement processing logic.
100
Daniel Estévez @destevez.net · 15/06/2026
New blog post: Tianwen-2 low data rate telemetry. This is a short post where I analyze a recording of low rate (4 kbaud) telemetry done on June 8 during Tianwen-2's arrival to asteroid Kamo'oalewa. destevez.net/2026/06/tian...
A screenshot of a GNU Radio GUI showing the spectrum and constellation of the signal.
0134
Daniel Estévez @destevez.net · 11/06/2026
Explicit formulas can be obtained by applying a Padé approximant to this degree 6 equation. This leads to designs that have a very small error in the loop bandwidth and have the correct damping. Read more: destevez.net/2026/06/pll-...
A plot that shows the loop bandwidth relative error in terms of the normalized loop bandwidth. The error tends to zero when the loop bandwidth tends to zero, and it is already around 0.01% for a loop bandwidth of 0.1.
000
Daniel Estévez @destevez.net · 11/06/2026
I showed that the loop coefficients for an order 2 loop with supercritical damping and phase/phase-rate update have explicit formulas that are obtained by solving a cubic equation. Now I treat the rate-only update case, showing that it leads to a degree 6 equation.
100
Daniel Estévez @destevez.net · 11/06/2026
New blog post: PLL coefficients for rate-only feedback. A couple years ago I wrote a post that extends a paper of Stephens and Thomas where they model PLLs in discrete time rather than transforming a continuous-time design.
101
Reposted by Daniel Estévez
AMSAT-DL @amsat-dl.org · 10/06/2026
Tianwen-2 at #Asteroid (469219) Kamoʻoalewa now transmitting at High Data Rate!! Receiving with the 20m #Bochum antenna on X-Band.
02410
Reposted by Daniel Estévez
Radiotelescoop Dwingeloo @radiotelescoop.bsky.social · 08/06/2026
We think a Tianwen-2 manoever happened early 7 June, as expected. With the Bochum telescope (@amsat-dl.org) and the Dwingeloo telescope, we observe that a) Tianwen-2 is close to the asteroid on the sky and b) the change in the line-of-sight velocity (Doppler) now almost matches that of the asteroid.
Sky plot showing the location where the Dwingeloo telescope measured the Tianwen-2 spacecraft, with a 0.1 degree field of view. On 7 June, the spacecraft is apparently within 0.1 degree of the asteroid.Frequency residual w.r.t. Kamo'oalewa, with an estimated base freq of 8428.201 MHz. On 4-6 June the residual is between 13800 and 14300 Hz, increasing. From 7 June, the line is almost flat, changing only a few Hz in hours. This shows the change in line-of-sight velocity is now close to that of Kamo'oalewa.
01710
Reposted by Daniel Estévez
AMSAT-DL @amsat-dl.org · 07/06/2026
Just a quick update on Tianwen-2. We saw several events before LOS- At 23:06:10 UTC the data rate was reduced, probably in preparation for some trajectory maneuvers.
192
Daniel Estévez @destevez.net · 03/06/2026
I saw it this morning. Quite interesting work!
000
Daniel Estévez @destevez.net · 28/05/2026
Second, @amsat-dl.org has decoded around 12 hours of telemetry. I run all this telemetry through the same analysis, showing that there are no relevant differences compared to the telemetry received by @radiotelescoop.bsky.social. Read more: destevez.net/2026/05/an-u...
destevez.net
An update about Tianwen-2 telemetry – Daniel Estévez
021
Daniel Estévez @destevez.net · 28/05/2026
New blog post: An update about Tianwen-2 telemetry. This is a quick update about Tianwen-2 telemetry decoding. First, I've figured out the format of the timestamps included in AOS frames. I explain what was confusing me yesterday.
A plot of space packet APIDs received over time by Bochum
140
Daniel Estévez @destevez.net · 27/05/2026
New blog post: Decoding Tianwen-2. I analyze the telemetry decoded from recent @radiotelescoop.bsky.social recordings of Tianwen-2. Compared to Tianwen-1, there is not much in the telemetry. There are not that many fields and most are mainly static. Read more: destevez.net/2026/05/deco...
A screenshot of a GNU Radio flowgraph GUI demodulating the signal from Tianwen-2, which has rather high SNR.A raster map of one of the decoded Space Packet APIDs. We can see a few fields changing values, but most of the contents are fairly static.
0112
Daniel Estévez @destevez.net · 13/05/2026
This is somewhat hidden to the programmer because the linker script tries to place the buffers in different banks when possible. However, it's relevant when there are more buffers than banks or when trying to access multiple elements from the same buffer, which is what I ran into.
000
Daniel Estévez @destevez.net · 13/05/2026
Each bank supports one simultaneous access, and every two banks are interleaved to form a 512-bit wide 16 KiB virtual bank. The two load units of the processor can perform simultaneous 512-bit loads only if these target different 16 KiB banks.
100
Daniel Estévez @destevez.net · 13/05/2026
More fun with the Ryzen AI NPU. I'm working on a more complex algorithm, and I'm seeing memory stalls that make me lose one cycle in each iteration. It turns out that the data memory of compute tiles is organized as 8x 256-bit wide 8 KiB memory banks.
A trace shown in the Perfetto UI that shows that the processor performs 8 cycles of vector instructions and then stalls waiting for memory for one cycle.A figure from the AIE-MLv2 documentation that shows the memory module organized as 8 banks.
110
Daniel Estévez @destevez.net · 08/05/2026
This post is an ideal self-contained introduction if you want to learn how NPUs work from a low-level perspective. Read more: destevez.net/2026/05/gett...
destevez.net
Getting peak TOPS on a Ryzen AI 7 350 NPU – Daniel Estévez
030
Daniel Estévez @destevez.net · 08/05/2026
Finally, I show how to use tracing to measure the performance of the NPU workload execution, and check that it matches the understanding we had obtained by analyzing the assembly code.
A screenshot of the Perfetto UI that shows an execution trace. We measure how many cycles it takes to execute each C++ kernel call.
110
Daniel Estévez @destevez.net · 08/05/2026
I explain how the IRON Python API is used to generate LLVM MLIR that defines how the NPU is set up, including the configuration of all the DMAs used for data movement. I go through the relevant sections of the MLIR code and explain how it is compiled to lower level objects.
110
Daniel Estévez @destevez.net · 08/05/2026
I implement a C++ kernel and show how it maps to assembly, and how to read the assembly to detect performance losses.
The assembly code for a C++ kernel that performs some 8x8 times 8x8 matrix multiplications. The architecture is VLIW and there are loop registers for hardware-controlled iteration. The main loop is just two instructions that are simultaneously loading data from memory with two load units and performing matrix multiply-accumulate with the vector unit.
110
Daniel Estévez @destevez.net · 08/05/2026
Then I explain how SIMD operations are fundamentally intended for matrix multiplication operations. For instance, an 8x8 times 8x8 matrix multiplication of int8 values can be done in a single SIMD instruction which performs 1024 integer operations in a single clock cycle.
100
Daniel Estévez @destevez.net · 08/05/2026
I start by giving an overview of the NPU hardware, explaining how it is organized as an array of compute, memory, and shimNOC tiles connected together mainly by an AXI-S interconnect for wide bandwidth data movement. I also explain the exposed-pipeline VLIW SIMD architecture.
A diagram that shows an array of compute tiles and their local memories. Data interconnections between the tiles are shown, including connection of each tile to an array-wide AXI-Stream interconnect, direct memory connections between adjacent tiles, and a cascade connection for accumulator values.
100
Daniel Estévez @destevez.net · 08/05/2026
New blog post: Getting peak TOPS on a Ryzen AI 7 350 NPU. This is an introduction to low-level programming on AMD NPUs using mlir-aie. I build an example that demonstrates 56 TOPS, very close to the max theoretical performance. These NPUs are identical to Xilinx AIE-MLv2 engines.
110
Daniel Estévez @destevez.net · 27/04/2026
I hadn't heard about keyoxide before, and this encouraged me to set up my profile. I agree that it's easy to do. keyoxide.org/aspe:keyoxid...
010
Daniel Estévez @destevez.net · 27/04/2026
This is to verify my Bluesky profile in Keyoxide: keyoxide.org/aspe%3Akeyox...
keyoxide.org
Daniel Estévez - Keyoxide
Modern and secure platform to manage a decentralized identity based on cryptographic keys
000
Daniel Estévez @destevez.net · 27/04/2026
aspe:keyoxide.org:MPXJS6MJRS3LENLWBNOOAMHOMY
100
Reposted by Daniel Estévez
AMSAT-DL @amsat-dl.org · 26/04/2026
Intruder Satellites on lower end of the 430 MHz UHF Amateur Radio band?
forum.amsat-dl.org
LORA Satellites on 430 MHz - AMSAT-DL Forum
Hello, Unidentified Russian satellites launched a few days ago alongside the Russian COSMOS-2600 military satellite: LoRa modulation at 430.630 MHz, 430.750 MHz and 430.910 MHz. Apparently quite st...
2135