Instructions to use replicate/paged-attention with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use replicate/paged-attention with Kernels:
# !pip install kernels from kernels import get_kernel kernel = get_kernel("replicate/paged-attention") - Notebooks
- Google Colab
- Kaggle
File size: 498 Bytes
132e594 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 | import torch
# Reference default values of atol and rtol are from
# https://github.com/pytorch/pytorch/blob/6d96beb6bec24d73ee3f080bac54d2104068f675/test/test_transformers.py#L67
default_atol = {torch.float16: 1e-3, torch.bfloat16: 1e-3, torch.float: 1e-5}
default_rtol = {torch.float16: 1e-3, torch.bfloat16: 1.6e-2, torch.float: 1.3e-6}
def get_default_atol(output) -> float:
return default_atol[output.dtype]
def get_default_rtol(output) -> float:
return default_rtol[output.dtype]
|