- Eliminate graph breaks in na.py and attention.py for full torch.compile support
* Replace cumsum-based tensor slicing with _tensor_split to avoid .item() calls
* Use torch.tensor_split with .long().cpu() for PyTorch API requirements
* Replace reshape-based averaging with split-stack-mean pattern
- Fix phase peak VRAM tracking to capture peaks during OOM retry cycles
- Standardize phase4 naming to 'postprocessing' for consistency
- Update RoPE docstrings for accuracy
- Update example workflows with icon