NOVA / ui /gpu.py
S-4-G-4-R's picture
Initial commit
20b15f3
Raw
History Blame Contribute Delete
2.4 kB
"""
ui/gpu.py
---------
ZeroGPU wiring — the only place in NOVA that touches a GPU.
HF's ZeroGPU hardware hands a Space a GPU *only* for the duration of a call to an
@spaces.GPU-decorated function, and refuses to boot at all if it can't find one
at import time ("No @spaces.GPU function detected during startup"). That single
constraint drives the whole design here.
NOVA is CPU-first and stays that way. Exactly one operation runs on the GPU: the
bulk embedding of a paper's chunks during vectorizing, which is by far the
slowest thing in the app (tens of seconds on CPU, a few on GPU). Everything else
— SPECTER reranking, query embedding, the cross-encoder — is pinned to CPU *on
purpose*, because those run outside any GPU window and a cuda-resident model
there would fail on first use. That's why the three backend call sites now take
an explicit `device` instead of auto-detecting.
Off ZeroGPU (local, or Spaces "CPU basic") @spaces.GPU is a transparent
passthrough and ON_ZEROGPU is False, so this module quietly degrades to plain
CPU work and nothing else in the app changes.
"""
import os
import spaces
# Set by the ZeroGPU runtime; `spaces.config` reads the same variable.
ON_ZEROGPU = os.getenv("SPACES_ZERO_GPU", "").lower() in ("1", "t", "true")
# The device to use *inside* a GPU window. Outside one, always "cpu".
GPU_DEVICE = "cuda" if ON_ZEROGPU else "cpu"
# Generous but bounded. The window has to cover PDF text extraction and chunking
# (CPU work that unavoidably happens inside build_vectorstore) plus the encode
# itself. A long paper on a cold cache is the worst case.
_VECTORIZE_SECONDS = 120
@spaces.GPU(duration=_VECTORIZE_SECONDS)
def vectorize_on_gpu(pdf_path: str) -> None:
"""Build and persist this paper's vectorstore with the embedder on GPU.
Returns None deliberately. ZeroGPU runs this in its own GPU worker, so a
Chroma handle created here would carry a cuda-resident embedding model back
to a caller that no longer holds the GPU — useless at best, a crash at worst.
What crosses the boundary is the *persisted vectorstore on disk*, which is
device-independent.
The caller then re-opens it on CPU, which costs nothing: build_vectorstore
short-circuits to a plain load as soon as the persist dir exists.
"""
from vectorizeer import build_vectorstore
build_vectorstore(pdf_path, device=GPU_DEVICE)