Spaces:
Running on Zero
Running on Zero
Deploy from 3dvalley spaces/material-maps
Browse files- .3dvalley +1 -0
- README.md +25 -7
- app.py +349 -0
- requirements.txt +15 -0
- upsampler_theme.py +54 -0
.3dvalley
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
material-maps
|
README.md
CHANGED
|
@@ -1,13 +1,31 @@
|
|
| 1 |
---
|
| 2 |
-
title: Material Maps
|
| 3 |
-
emoji:
|
| 4 |
-
colorFrom:
|
| 5 |
-
colorTo:
|
| 6 |
sdk: gradio
|
| 7 |
-
sdk_version: 6.
|
| 8 |
-
python_version:
|
| 9 |
app_file: app.py
|
| 10 |
pinned: false
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
---
|
| 12 |
|
| 13 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
title: "Material Maps - Normal, Height, Roughness from One Image"
|
| 3 |
+
emoji: 🧱
|
| 4 |
+
colorFrom: indigo
|
| 5 |
+
colorTo: purple
|
| 6 |
sdk: gradio
|
| 7 |
+
sdk_version: 6.1.0
|
| 8 |
+
python_version: "3.12"
|
| 9 |
app_file: app.py
|
| 10 |
pinned: false
|
| 11 |
+
models:
|
| 12 |
+
- InvokeAI/pbr-material-maps
|
| 13 |
+
- jingheya/lotus-normal-g-v1-1
|
| 14 |
+
- openai/clip-vit-base-patch32
|
| 15 |
+
license: apache-2.0
|
| 16 |
+
short_description: "PBR normal, height, roughness, metallic maps from a texture."
|
| 17 |
---
|
| 18 |
|
| 19 |
+
# Material Maps - Normal, Height and Roughness from One Image
|
| 20 |
+
|
| 21 |
+
Turn one picture of a surface, a texture tile or a photo, into the maps a PBR renderer needs: a tangent-space normal map (OpenGL, or DirectX on request), a 16-bit height map, roughness and metallic. The maps come back at the picture's size, up to 1024 pixels, and a seamless picture gives seamless maps: every network pads by wrapping around the edges.
|
| 22 |
+
|
| 23 |
+
Height is not guessed from brightness. Nets trained on texture sets (the Material Map Generator ESRGAN models) read painted bricks as raised and white mortar as sunk, which a brightness-based height map gets backwards. Lotus-G, a surface-normal diffusion model, adds the broad shape of stones and bricks, and CLIP names the material to set how rough it is and whether it is metal.
|
| 24 |
+
|
| 25 |
+
API endpoint, one GPU call, about a second on the GPU:
|
| 26 |
+
|
| 27 |
+
- `/material_maps(image, directx=False)` returns `normal.png`, `height.png` (16-bit), `roughness.png`, `metallic.png` and a JSON note (`material`, `classes`, `roughness_level`, `metal`, `seconds`).
|
| 28 |
+
|
| 29 |
+
## Free on 3D Valley
|
| 30 |
+
|
| 31 |
+
Used by the texture generator and the normal map tool on [3D Valley](https://3dvalley.com). Built by [Upsampler](https://upsampler.com), which also offers AI image generation, editing, upscaling, and enhancement tools.
|
app.py
ADDED
|
@@ -0,0 +1,349 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""
|
| 2 |
+
Material maps for 3dvalley.com: one picture of a surface (a texture tile or a
|
| 3 |
+
photo) in, the maps a PBR renderer needs out, all tiling when the picture
|
| 4 |
+
tiles. One API endpoint for headless callers (the site's browser client),
|
| 5 |
+
plus a small demo UI.
|
| 6 |
+
|
| 7 |
+
How a run goes, all on the GPU in one call:
|
| 8 |
+
- Two small ESRGAN nets trained on texture sets (Joey Ballentine's Material
|
| 9 |
+
Map Generator, Apache-2.0): one gives a tangent-space normal map, the other
|
| 10 |
+
displacement and roughness. They are what reads a painted brick as raised
|
| 11 |
+
and its white mortar as sunk, which brightness alone gets backwards.
|
| 12 |
+
- Lotus-G normal (Apache-2.0, Stable Diffusion 2 fine-tuned for surface
|
| 13 |
+
normals, one step) gives the broad shape: the rounded top of a cobble, the
|
| 14 |
+
bevel of a brick. Its low frequencies and the ESRGAN detail are merged in
|
| 15 |
+
slope space with the slopes of the displacement map, so normal and height
|
| 16 |
+
agree.
|
| 17 |
+
- CLIP (MIT) names the material class, which sets the roughness level and
|
| 18 |
+
whether anything is metal; the ESRGAN roughness adds the variation.
|
| 19 |
+
|
| 20 |
+
Every convolution pads circularly (the ESRGAN input is wrapped, the Lotus UNet
|
| 21 |
+
and VAE have their padding mode switched), and every filter wraps, so a
|
| 22 |
+
seamless picture gives seamless maps.
|
| 23 |
+
"""
|
| 24 |
+
|
| 25 |
+
import os
|
| 26 |
+
import tempfile
|
| 27 |
+
import time
|
| 28 |
+
|
| 29 |
+
import spaces
|
| 30 |
+
|
| 31 |
+
os.environ["GRADIO_TEMP_DIR"] = os.path.join(tempfile.gettempdir(), "gradio")
|
| 32 |
+
os.makedirs(os.environ["GRADIO_TEMP_DIR"], exist_ok=True)
|
| 33 |
+
|
| 34 |
+
import gradio as gr
|
| 35 |
+
import numpy as np
|
| 36 |
+
import torch
|
| 37 |
+
import torch.nn.functional as F
|
| 38 |
+
from diffusers import AutoencoderKL, UNet2DConditionModel
|
| 39 |
+
from huggingface_hub import hf_hub_download
|
| 40 |
+
from PIL import Image
|
| 41 |
+
from spandrel import ModelLoader
|
| 42 |
+
from transformers import CLIPModel, CLIPProcessor, CLIPTextModel, CLIPTokenizer
|
| 43 |
+
|
| 44 |
+
from upsampler_theme import UPSAMPLER_CSS, UPSAMPLER_THEME, footer_html, header_html
|
| 45 |
+
|
| 46 |
+
DEVICE = "cuda"
|
| 47 |
+
MAX_SIDE = 1024
|
| 48 |
+
LOTUS_SIDE = 768 # Lotus-G is Stable Diffusion 2 base: 512 to 768 is home.
|
| 49 |
+
MAPS_REPO = "InvokeAI/pbr-material-maps"
|
| 50 |
+
MAPS_REVISION = "b7ca9ebc6e14688a69d41872d2b9c80ea453e8f0"
|
| 51 |
+
LOTUS_REPO = "jingheya/lotus-normal-g-v1-1"
|
| 52 |
+
CLIP_REPO = "openai/clip-vit-base-patch32"
|
| 53 |
+
|
| 54 |
+
|
| 55 |
+
def _log(*parts):
|
| 56 |
+
print("[maps]", *parts, flush=True)
|
| 57 |
+
|
| 58 |
+
|
| 59 |
+
def _circular(module: torch.nn.Module) -> None:
|
| 60 |
+
for m in module.modules():
|
| 61 |
+
if isinstance(m, torch.nn.Conv2d) and m.padding not in (0, (0, 0)):
|
| 62 |
+
m.padding_mode = "circular"
|
| 63 |
+
|
| 64 |
+
|
| 65 |
+
def _esrgan(name: str) -> torch.nn.Module:
|
| 66 |
+
path = hf_hub_download(MAPS_REPO, name, revision=MAPS_REVISION)
|
| 67 |
+
return ModelLoader().load_from_file(path).model.eval().half().to(DEVICE)
|
| 68 |
+
|
| 69 |
+
|
| 70 |
+
normal_net = _esrgan("normal_map_generator.safetensors")
|
| 71 |
+
franken_net = _esrgan("franken_map_generator.safetensors")
|
| 72 |
+
|
| 73 |
+
lotus_unet = UNet2DConditionModel.from_pretrained(LOTUS_REPO, subfolder="unet", torch_dtype=torch.float16).to(DEVICE)
|
| 74 |
+
lotus_vae = AutoencoderKL.from_pretrained(LOTUS_REPO, subfolder="vae", torch_dtype=torch.float16).to(DEVICE)
|
| 75 |
+
_circular(lotus_unet)
|
| 76 |
+
_circular(lotus_vae)
|
| 77 |
+
# Lotus runs with an empty prompt: encode it once and drop the text encoder.
|
| 78 |
+
with torch.no_grad():
|
| 79 |
+
_tok = CLIPTokenizer.from_pretrained(LOTUS_REPO, subfolder="tokenizer")
|
| 80 |
+
_enc = CLIPTextModel.from_pretrained(LOTUS_REPO, subfolder="text_encoder")
|
| 81 |
+
_ids = _tok([""], padding="max_length", max_length=_tok.model_max_length, return_tensors="pt").input_ids
|
| 82 |
+
EMPTY_PROMPT = _enc(_ids)[0].half().to(DEVICE)
|
| 83 |
+
del _tok, _enc
|
| 84 |
+
# The task embedding that selects the normal head (see Lotus's infer.py).
|
| 85 |
+
_task = torch.tensor([[1.0, 0.0]])
|
| 86 |
+
TASK_EMB = torch.cat([torch.sin(_task), torch.cos(_task)], dim=-1).half().to(DEVICE)
|
| 87 |
+
|
| 88 |
+
clip_model = CLIPModel.from_pretrained(CLIP_REPO, torch_dtype=torch.float16).eval().to(DEVICE)
|
| 89 |
+
clip_processor = CLIPProcessor.from_pretrained(CLIP_REPO)
|
| 90 |
+
|
| 91 |
+
# (label, words for CLIP, roughness level, metal): "metal" is bare metal all
|
| 92 |
+
# over, "rust" is metal only where grey steel shows through.
|
| 93 |
+
CLASSES = [
|
| 94 |
+
("brick", "a brick wall texture", 0.85, None),
|
| 95 |
+
("stone", "a cobblestone or stone paving texture", 0.8, None),
|
| 96 |
+
("rock", "a rough natural rock texture", 0.85, None),
|
| 97 |
+
("concrete", "a concrete or plaster wall texture", 0.9, None),
|
| 98 |
+
("asphalt", "an asphalt road texture", 0.9, None),
|
| 99 |
+
("wood", "a wooden planks texture", 0.7, None),
|
| 100 |
+
("varnished wood", "a varnished polished wood floor texture", 0.35, None),
|
| 101 |
+
("bark", "a tree bark texture", 0.9, None),
|
| 102 |
+
("ground", "a dirt, soil, mud or sand ground texture", 0.95, None),
|
| 103 |
+
("vegetation", "a grass, moss or leaves texture", 0.8, None),
|
| 104 |
+
("marble", "a polished marble texture", 0.2, None),
|
| 105 |
+
("tiles", "a glazed ceramic tiles texture", 0.25, None),
|
| 106 |
+
("fabric", "a fabric, cloth or carpet texture", 0.9, None),
|
| 107 |
+
("leather", "a leather texture", 0.6, None),
|
| 108 |
+
("plastic", "a plastic surface texture", 0.4, None),
|
| 109 |
+
("painted metal", "a painted metal surface texture", 0.5, None),
|
| 110 |
+
("rusted metal", "a rusty corroded metal texture", 0.8, "rust"),
|
| 111 |
+
("brushed metal", "a brushed steel or aluminium metal texture", 0.35, "metal"),
|
| 112 |
+
("polished metal", "a shiny polished metal, chrome, gold or copper texture", 0.15, "metal"),
|
| 113 |
+
("snow", "a snow or ice texture", 0.3, None),
|
| 114 |
+
]
|
| 115 |
+
with torch.no_grad():
|
| 116 |
+
_t = clip_processor(text=[c[1] for c in CLASSES], return_tensors="pt", padding=True).to(DEVICE)
|
| 117 |
+
CLASS_EMB = F.normalize(clip_model.get_text_features(**_t).float(), dim=-1)
|
| 118 |
+
_log("models ready")
|
| 119 |
+
|
| 120 |
+
|
| 121 |
+
# --- plain-array helpers, all wrapping at the edges -------------------------
|
| 122 |
+
|
| 123 |
+
def _blur(a: torch.Tensor, sigma: float) -> torch.Tensor:
|
| 124 |
+
"""Separable Gaussian on an (H, W) tensor, wrapping around the edges."""
|
| 125 |
+
if sigma <= 0:
|
| 126 |
+
return a
|
| 127 |
+
radius = max(1, int(3 * sigma))
|
| 128 |
+
x = torch.arange(-radius, radius + 1, device=a.device, dtype=a.dtype)
|
| 129 |
+
k = torch.exp(-(x**2) / (2 * sigma**2))
|
| 130 |
+
k = k / k.sum()
|
| 131 |
+
out = F.pad(a[None, None], (radius, radius, 0, 0), mode="circular")
|
| 132 |
+
out = F.conv2d(out, k.view(1, 1, 1, -1))
|
| 133 |
+
out = F.pad(out, (0, 0, radius, radius), mode="circular")
|
| 134 |
+
return F.conv2d(out, k.view(1, 1, -1, 1))[0, 0]
|
| 135 |
+
|
| 136 |
+
|
| 137 |
+
def _resize_wrap(x: torch.Tensor, size: tuple[int, int]) -> torch.Tensor:
|
| 138 |
+
"""Resize (N, C, H, W) so the result still tiles: pad by wrapping, scale, crop."""
|
| 139 |
+
h, w = x.shape[-2:]
|
| 140 |
+
if (h, w) == size:
|
| 141 |
+
return x
|
| 142 |
+
pad = 4
|
| 143 |
+
big = F.pad(x, (pad, pad, pad, pad), mode="circular")
|
| 144 |
+
sy, sx = size[0] / h, size[1] / w
|
| 145 |
+
out = F.interpolate(big, size=(round((h + 2 * pad) * sy), round((w + 2 * pad) * sx)), mode="bicubic", align_corners=False)
|
| 146 |
+
oy, ox = round(pad * sy), round(pad * sx)
|
| 147 |
+
return out[..., oy : oy + size[0], ox : ox + size[1]]
|
| 148 |
+
|
| 149 |
+
|
| 150 |
+
def _slopes(n: torch.Tensor) -> tuple[torch.Tensor, torch.Tensor]:
|
| 151 |
+
"""(3, H, W) normals, x right, y up → slopes dh/dx and dh/dy_up, tilt removed.
|
| 152 |
+
A normal is (-dh/dx, -dh/dy, 1) normalised."""
|
| 153 |
+
nz = n[2].clamp(min=0.2)
|
| 154 |
+
p, q = -n[0] / nz, -n[1] / nz
|
| 155 |
+
return p - p.mean(), q - q.mean()
|
| 156 |
+
|
| 157 |
+
|
| 158 |
+
def _grad(h: torch.Tensor) -> tuple[torch.Tensor, torch.Tensor]:
|
| 159 |
+
"""Central differences with wrap: dh/dx and dh/dy_up (rows run down)."""
|
| 160 |
+
gx = (torch.roll(h, -1, 1) - torch.roll(h, 1, 1)) / 2
|
| 161 |
+
gy = (torch.roll(h, 1, 0) - torch.roll(h, -1, 0)) / 2
|
| 162 |
+
return gx, gy
|
| 163 |
+
|
| 164 |
+
|
| 165 |
+
def _stretch(a: torch.Tensor, lo: float = 0.005, hi: float = 0.995) -> torch.Tensor:
|
| 166 |
+
flat = a.flatten()
|
| 167 |
+
if flat.numel() > 1_000_000:
|
| 168 |
+
flat = flat[:: flat.numel() // 1_000_000 + 1]
|
| 169 |
+
a_lo, a_hi = torch.quantile(flat, lo), torch.quantile(flat, hi)
|
| 170 |
+
return ((a - a_lo) / (a_hi - a_lo).clamp(min=1e-6)).clamp(0, 1)
|
| 171 |
+
|
| 172 |
+
|
| 173 |
+
def _smoothstep(e0: float, e1: float, x: torch.Tensor) -> torch.Tensor:
|
| 174 |
+
t = ((x - e0) / (e1 - e0)).clamp(0, 1)
|
| 175 |
+
return t * t * (3 - 2 * t)
|
| 176 |
+
|
| 177 |
+
|
| 178 |
+
# --- the models --------------------------------------------------------------
|
| 179 |
+
|
| 180 |
+
def _run_esrgan(net: torch.nn.Module, rgb: torch.Tensor) -> torch.Tensor:
|
| 181 |
+
"""(1, 3, H, W) in [0, 1] → (3, H, W) in [0, 1]; wrapped so the edges tile."""
|
| 182 |
+
pad = 32
|
| 183 |
+
x = F.pad(rgb, (pad, pad, pad, pad), mode="circular").half()
|
| 184 |
+
return net(x)[0, :, pad:-pad, pad:-pad].float().clamp(0, 1)
|
| 185 |
+
|
| 186 |
+
|
| 187 |
+
def _run_lotus(rgb: torch.Tensor) -> torch.Tensor:
|
| 188 |
+
"""(1, 3, H, W) in [0, 1] → (3, H, W) unit normals, x right, y up, z out."""
|
| 189 |
+
h, w = rgb.shape[-2:]
|
| 190 |
+
scale = min(1.0, LOTUS_SIDE / max(h, w))
|
| 191 |
+
size = (max(64, round(h * scale / 64) * 64), max(64, round(w * scale / 64) * 64))
|
| 192 |
+
x = _resize_wrap(rgb, size) * 2 - 1
|
| 193 |
+
latents = lotus_vae.encode(x.half()).latent_dist.mode() * lotus_vae.config.scaling_factor
|
| 194 |
+
noise = torch.randn(latents.shape, generator=torch.Generator(DEVICE).manual_seed(0), device=DEVICE, dtype=latents.dtype)
|
| 195 |
+
x0 = lotus_unet(
|
| 196 |
+
torch.cat([latents, noise], dim=1),
|
| 197 |
+
torch.tensor([999], device=DEVICE),
|
| 198 |
+
encoder_hidden_states=EMPTY_PROMPT,
|
| 199 |
+
class_labels=TASK_EMB,
|
| 200 |
+
return_dict=False,
|
| 201 |
+
)[0]
|
| 202 |
+
decoded = lotus_vae.decode(x0 / lotus_vae.config.scaling_factor, return_dict=False)[0].float().clamp(-1, 1)
|
| 203 |
+
n = _resize_wrap(decoded, (h, w))[0]
|
| 204 |
+
return n / n.norm(dim=0, keepdim=True).clamp(min=1e-6)
|
| 205 |
+
|
| 206 |
+
|
| 207 |
+
def _classify(image: Image.Image) -> tuple[torch.Tensor, list[tuple[str, float]]]:
|
| 208 |
+
inputs = clip_processor(images=image, return_tensors="pt").to(DEVICE)
|
| 209 |
+
emb = F.normalize(clip_model.get_image_features(pixel_values=inputs.pixel_values.half()).float(), dim=-1)
|
| 210 |
+
probs = (100 * emb @ CLASS_EMB.T).softmax(dim=-1)[0]
|
| 211 |
+
order = probs.argsort(descending=True)[:3].tolist()
|
| 212 |
+
return probs, [(CLASSES[i][0], round(float(probs[i]), 3)) for i in order]
|
| 213 |
+
|
| 214 |
+
|
| 215 |
+
@spaces.GPU(duration=20)
|
| 216 |
+
@torch.no_grad()
|
| 217 |
+
def _maps(image: Image.Image):
|
| 218 |
+
t0 = time.time()
|
| 219 |
+
w, h = image.size
|
| 220 |
+
rgb = torch.from_numpy(np.asarray(image, np.float32) / 255).permute(2, 0, 1)[None].to(DEVICE)
|
| 221 |
+
s = max(w, h) / 768 # filter sizes were tuned at 768 px
|
| 222 |
+
|
| 223 |
+
es_normal = _run_esrgan(normal_net, rgb) * 2 - 1
|
| 224 |
+
franken = _run_esrgan(franken_net, rgb)
|
| 225 |
+
lotus = _run_lotus(rgb)
|
| 226 |
+
probs, top = _classify(image)
|
| 227 |
+
_log(f"models {time.time() - t0:.2f}s", top)
|
| 228 |
+
|
| 229 |
+
# Height: the texture-trained displacement. Normal: the displacement's
|
| 230 |
+
# slopes, plus Lotus's broad shape and the ESRGAN normal's fine detail.
|
| 231 |
+
height = _stretch(franken[2])
|
| 232 |
+
pd, qd = _grad(height)
|
| 233 |
+
pd, qd = pd * max(w, h) / 40, qd * max(w, h) / 40
|
| 234 |
+
pl, ql = _slopes(lotus)
|
| 235 |
+
pe, qe = _slopes(es_normal)
|
| 236 |
+
p = (_blur(pl, 2 * s) + pe - _blur(pe, 3 * s) + pd) / 2
|
| 237 |
+
q = (_blur(ql, 2 * s) + qe - _blur(qe, 3 * s) + qd) / 2
|
| 238 |
+
normal = torch.stack([-p, -q, torch.ones_like(p)])
|
| 239 |
+
normal = normal / normal.norm(dim=0, keepdim=True)
|
| 240 |
+
|
| 241 |
+
# Roughness: the class sets the level, the ESRGAN map the variation.
|
| 242 |
+
level = sum(float(probs[i]) * c[2] for i, c in enumerate(CLASSES))
|
| 243 |
+
rough = franken[1]
|
| 244 |
+
roughness = (level + (rough - rough.median()) * 1.2).clamp(0.04, 1)
|
| 245 |
+
|
| 246 |
+
# Metallic: bare metal is metal all over; rusted metal only where grey
|
| 247 |
+
# steel shows (low saturation). Everything else is not metal.
|
| 248 |
+
metal = sum(float(probs[i]) for i, c in enumerate(CLASSES) if c[3] == "metal")
|
| 249 |
+
rust = sum(float(probs[i]) for i, c in enumerate(CLASSES) if c[3] == "rust")
|
| 250 |
+
mx, mn = rgb[0].max(dim=0).values, rgb[0].min(dim=0).values
|
| 251 |
+
saturation = (mx - mn) / mx.clamp(min=1e-3)
|
| 252 |
+
bare = _smoothstep(0.35, 0.15, saturation)
|
| 253 |
+
metallic = _smoothstep(0.35, 0.65, metal + rust * bare)
|
| 254 |
+
roughness = roughness - metallic * 0.15
|
| 255 |
+
|
| 256 |
+
info = {
|
| 257 |
+
"material": top[0][0],
|
| 258 |
+
"classes": [{"label": label, "p": p_} for label, p_ in top],
|
| 259 |
+
"roughness_level": round(level, 3),
|
| 260 |
+
"metal": round(metal, 3),
|
| 261 |
+
"gpu_seconds": round(time.time() - t0, 2),
|
| 262 |
+
}
|
| 263 |
+
out = (
|
| 264 |
+
(normal.permute(1, 2, 0) * 0.5 + 0.5).clamp(0, 1).cpu().numpy(),
|
| 265 |
+
height.cpu().numpy(),
|
| 266 |
+
roughness.clamp(0, 1).cpu().numpy(),
|
| 267 |
+
metallic.clamp(0, 1).cpu().numpy(),
|
| 268 |
+
)
|
| 269 |
+
torch.cuda.empty_cache()
|
| 270 |
+
return out, info
|
| 271 |
+
|
| 272 |
+
|
| 273 |
+
def _save_png(array: np.ndarray, stem: str, bits: int = 8) -> str:
|
| 274 |
+
path = os.path.join(os.environ["GRADIO_TEMP_DIR"], f"{stem}-{time.time_ns()}.png")
|
| 275 |
+
if bits == 16:
|
| 276 |
+
Image.fromarray((array * 65535).round().astype(np.uint16)).save(path)
|
| 277 |
+
else:
|
| 278 |
+
Image.fromarray((array * 255).round().astype(np.uint8)).save(path, optimize=False, compress_level=6)
|
| 279 |
+
return path
|
| 280 |
+
|
| 281 |
+
|
| 282 |
+
def material_maps(image, directx: bool = False):
|
| 283 |
+
"""A picture of a surface → normal (OpenGL unless `directx`), height
|
| 284 |
+
(16-bit), roughness and metallic PNGs at its size (capped at 1024 px),
|
| 285 |
+
and a small JSON note of what the surface was taken for."""
|
| 286 |
+
if image is None:
|
| 287 |
+
raise gr.Error("Upload a picture of a surface.")
|
| 288 |
+
if not isinstance(image, Image.Image):
|
| 289 |
+
image = Image.open(image)
|
| 290 |
+
image = image.convert("RGB")
|
| 291 |
+
if max(image.size) > MAX_SIDE:
|
| 292 |
+
scale = MAX_SIDE / max(image.size)
|
| 293 |
+
image = image.resize((max(8, round(image.width * scale)), max(8, round(image.height * scale))), Image.LANCZOS)
|
| 294 |
+
t0 = time.time()
|
| 295 |
+
(normal, height, roughness, metallic), info = _maps(image)
|
| 296 |
+
if directx:
|
| 297 |
+
normal = normal.copy()
|
| 298 |
+
normal[..., 1] = 1 - normal[..., 1]
|
| 299 |
+
info["convention"] = "directx" if directx else "opengl"
|
| 300 |
+
info["size"] = [image.width, image.height]
|
| 301 |
+
files = (
|
| 302 |
+
_save_png(normal, "normal"),
|
| 303 |
+
_save_png(height, "height", bits=16),
|
| 304 |
+
_save_png(roughness, "roughness"),
|
| 305 |
+
_save_png(metallic, "metallic"),
|
| 306 |
+
)
|
| 307 |
+
info["seconds"] = round(time.time() - t0, 2)
|
| 308 |
+
_log("done", info)
|
| 309 |
+
return (*files, info)
|
| 310 |
+
|
| 311 |
+
|
| 312 |
+
with gr.Blocks(title="Material Maps - Normal, Height and Roughness from One Image") as demo:
|
| 313 |
+
gr.HTML(header_html(
|
| 314 |
+
"Material Maps",
|
| 315 |
+
"Normal, height, roughness and metallic maps from one picture of a surface. Seamless in, seamless out.",
|
| 316 |
+
))
|
| 317 |
+
with gr.Row(equal_height=False):
|
| 318 |
+
with gr.Column():
|
| 319 |
+
src = gr.Image(type="pil", image_mode="RGB", label="Texture or photo of a surface", height=360)
|
| 320 |
+
directx = gr.Checkbox(value=False, label="DirectX normal map (green down, for Unreal)")
|
| 321 |
+
btn = gr.Button("Make Maps", variant="primary")
|
| 322 |
+
with gr.Column():
|
| 323 |
+
with gr.Row():
|
| 324 |
+
out_normal = gr.Image(type="filepath", label="Normal", height=200)
|
| 325 |
+
out_height = gr.Image(type="filepath", label="Height (16-bit)", height=200)
|
| 326 |
+
with gr.Row():
|
| 327 |
+
out_rough = gr.Image(type="filepath", label="Roughness", height=200)
|
| 328 |
+
out_metal = gr.Image(type="filepath", label="Metallic", height=200)
|
| 329 |
+
out_info = gr.JSON(label="Surface")
|
| 330 |
+
btn.click(
|
| 331 |
+
material_maps,
|
| 332 |
+
inputs=[src, directx],
|
| 333 |
+
outputs=[out_normal, out_height, out_rough, out_metal, out_info],
|
| 334 |
+
api_name="material_maps",
|
| 335 |
+
)
|
| 336 |
+
gr.HTML(footer_html(
|
| 337 |
+
"Turn a texture or a photo of a surface into a PBR material: a tangent-space normal map, a 16-bit "
|
| 338 |
+
"height (displacement) map, roughness and metallic, at the picture's size up to 1024 pixels. Nets "
|
| 339 |
+
"trained on texture sets read painted bricks and stones the right way round, a surface-normal "
|
| 340 |
+
"diffusion model adds the broad shape, and every step wraps at the edges so seamless textures stay "
|
| 341 |
+
"seamless. Ready for Blender, Unity, Unreal, three.js and glTF.",
|
| 342 |
+
"https://upsampler.com",
|
| 343 |
+
"Upsampler",
|
| 344 |
+
))
|
| 345 |
+
|
| 346 |
+
if __name__ == "__main__":
|
| 347 |
+
demo.queue(default_concurrency_limit=2).launch(
|
| 348 |
+
theme=UPSAMPLER_THEME, css=UPSAMPLER_CSS, ssr_mode=False, show_error=True
|
| 349 |
+
)
|
requirements.txt
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# The torch/CUDA stack spaces/trellis-2 and spaces/hunyuan3d-paint run on
|
| 2 |
+
# ZeroGPU (Python 3.12, torch 2.11, CUDA 13).
|
| 3 |
+
--extra-index-url https://download.pytorch.org/whl/cu130
|
| 4 |
+
|
| 5 |
+
torch==2.11.0
|
| 6 |
+
torchvision==0.26.0
|
| 7 |
+
diffusers==0.35.2
|
| 8 |
+
transformers==4.57.3
|
| 9 |
+
accelerate==1.10.1
|
| 10 |
+
safetensors==0.6.2
|
| 11 |
+
spandrel==0.4.2
|
| 12 |
+
numpy==2.2.6
|
| 13 |
+
pillow==12.0.0
|
| 14 |
+
# `spaces` must NOT be pinned (HF injects its own) and gradio is set by the
|
| 15 |
+
# README sdk_version, so neither belongs here.
|
upsampler_theme.py
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Shared Upsampler look-and-feel for every Upsampler/* HF Space.
|
| 2 |
+
|
| 3 |
+
`create_and_push.py` uploads this file alongside each Space's app.py, so every
|
| 4 |
+
Space imports the exact same theme, CSS, header, and footer. Keep this the
|
| 5 |
+
single source of truth (v3 recipe): Soft indigo/purple theme, gradient primary
|
| 6 |
+
button, Gradio's own footer hidden, 1000px max width, minimal header/footer.
|
| 7 |
+
"""
|
| 8 |
+
|
| 9 |
+
import gradio as gr
|
| 10 |
+
|
| 11 |
+
UPSAMPLER_THEME = gr.themes.Soft(
|
| 12 |
+
primary_hue=gr.themes.colors.indigo,
|
| 13 |
+
secondary_hue=gr.themes.colors.purple,
|
| 14 |
+
neutral_hue=gr.themes.colors.slate,
|
| 15 |
+
font=[gr.themes.GoogleFont("Inter"), "system-ui", "sans-serif"],
|
| 16 |
+
).set(
|
| 17 |
+
button_primary_background_fill="linear-gradient(90deg, #6366f1 0%, #a855f7 100%)",
|
| 18 |
+
button_primary_background_fill_hover="linear-gradient(90deg, #4f46e5 0%, #9333ea 100%)",
|
| 19 |
+
button_primary_text_color="#ffffff",
|
| 20 |
+
button_primary_border_color="*primary_500",
|
| 21 |
+
)
|
| 22 |
+
|
| 23 |
+
# Hide Gradio's built-in footer and keep the app narrow and centered.
|
| 24 |
+
UPSAMPLER_CSS = """
|
| 25 |
+
footer { display: none !important; }
|
| 26 |
+
.gradio-container { max-width: 1000px !important; margin: 0 auto !important; }
|
| 27 |
+
#usp-header h1 { font-size: 1.7rem; font-weight: 700; margin: 0 0 .25rem; }
|
| 28 |
+
#usp-header p { opacity: .6; margin: 0; }
|
| 29 |
+
#usp-footer { opacity: .5; font-size: .85rem; margin-top: 1.25rem; }
|
| 30 |
+
#usp-footer a { text-decoration: none; }
|
| 31 |
+
"""
|
| 32 |
+
|
| 33 |
+
|
| 34 |
+
def header_html(title: str, subtitle: str) -> str:
|
| 35 |
+
return f"""<div id="usp-header">
|
| 36 |
+
<h1>{title}</h1>
|
| 37 |
+
<p>{subtitle}</p>
|
| 38 |
+
</div>"""
|
| 39 |
+
|
| 40 |
+
|
| 41 |
+
_LINK_STYLE = "color:#8b7cf6;font-weight:600;text-decoration:none"
|
| 42 |
+
|
| 43 |
+
|
| 44 |
+
def footer_html(description: str, tool_url: str, tool_anchor: str) -> str:
|
| 45 |
+
"""SEO footer (v4 recipe): a short paragraph describing what the model
|
| 46 |
+
does (unique per Space, keyword-bearing) plus the Upsampler attribution
|
| 47 |
+
with a deep link to the matching /free-* tool on upsampler.com. No model
|
| 48 |
+
credit/license line (the README frontmatter carries the license). Spaces
|
| 49 |
+
target model-name queries; the site pages keep the intent queries
|
| 50 |
+
("free X no signup"), so the two never compete."""
|
| 51 |
+
return f"""<div id="usp-footer">
|
| 52 |
+
<p style="margin:0 0 10px">{description}</p>
|
| 53 |
+
<p style="margin:0">Maintained by <a href="https://upsampler.com" target="_blank" rel="noopener" style="{_LINK_STYLE}">Upsampler</a>. Check out the <a href="{tool_url}" target="_blank" rel="noopener" style="{_LINK_STYLE}">{tool_anchor}</a>, no sign-up required.</p>
|
| 54 |
+
</div>"""
|