Posted by hsfzxjy 2 days ago
Obviously I am mostly just driving the DPU, as the SHAVE cores are just not beefy enough, but maybe this will be useful at some point, at least to gain more insight into the architecture...
As for the interesting cost variance for reshape you've mentioned, I guess surrounding context is the cause. The NPU compiler might adjust the data layout to match successive OPs' requirements. But, yes, that's annoying~
Studying DPU is not my current priority, but I may dig it up in the future to understand its invocation descriptor, and hopefully to discover more interesting stuff. I hope this project ends up being useful to you!
1. Implement uncommon NN operators.
2. Implement a high-precision (FP32, etc.) version of existing OpenVINO operators for numerically sensitive usage.
3. Implement a mega kernel that fuse several small kernels together to reduce SHAVE invocations, and potentially improve performance.
But most importantly, it gives you more control over the hardware you own. That control IMO should have been yours from day one you bought the machine.
plonk.
But based on my understanding, playing this on Linux or newer NPU generations might be possible. I have recorded my hypotheses in [0]. Verifying the hypotheses would require additional utility tools, which I am stilling crafting.
[0]: https://github.com/hsfzxjy/npunlock/blob/master/wiki/PORTING...