Direct Feedback Alignment and its variants fail to learn hidden hierarchies, even with near-optimal feedback weights.
The reason: DFA ignores input-specific masking - the derivative of forward activations.
In controlled experiments, we demonstrate its necessity in approximating backpropagation.