What vkML does today, counted from the tree at build time. Every number on this page is generated; none of it is typed.
Operators by category
| Category | Operators | On Vulkan |
|---|---|---|
Creation | 11 | 11 |
Element-wise | 22 | 22 |
Arithmetic | 7 | 7 |
Comparison | 7 | 7 |
Reduction | 7 | 6 |
Shape & indexing | 9 | 9 |
Linear algebra & NN | 10 | 10 |
Losses | 5 | 5 |
Autograd & execution | 5 | 5 |
Serialization | 4 | 4 |
Devices & introspection | 22 | 22 |
Explaining what the engine chose | 5 | 5 |
Data types
Five, on both backends: float32, float16,
int32, int64 and bool. f32→f16
narrowing is done in software rather than left to OpFConvert, whose
rounding mode SPIR-V leaves implementation-defined — so a narrowed value is
the same bits on every driver.
Devices
Any device exposing Vulkan 1.3 compute, plus the CPU backend. The CPU one is the correctness oracle the GPU is checked against, not a fallback: it is always built, always available, and every Vulkan result is compared against it.
Autograd
48 of 67 graph operations carry a gradient rule. The remaining 19 are listed on Current limitations with the reason for each — most of them are leaves or comparisons, where a gradient is not a thing that exists.
Serialization
Checkpoints are a zip of NumPy arrays plus a JSON manifest, loaded with
allow_pickle=False, written to a temporary file and renamed so an
interrupted save cannot leave a half-written checkpoint in place. A PyTorch
state_dict loads without translation.