# Deployment Profiling This specification defines how Harbor-HF chooses benchmark concurrency for one exact serving deployment. A selected profile is one immutable JSON document stored with its raw measurements in the private artifact Bucket. The general measurement method is engine-independent. Harbor-HF applies it to remote Inference Endpoints and Inference Providers; it never loads a model or runs benchmark tasks on the operator machine. ## Minimal Profile ```json { "schema_version": "harbor-hf/serving-profile/v1", "profile_id": "ornith-35b-q4-h200-64k-20260718", "created_at": "2026-07-18T01:00:00Z", "identity": { "model_sha256": "sha256:0000000000000000000000000000000000000000000000000000000000000000", "deployment_sha256": "sha256:0000000000000000000000000000000000000000000000000000000000000000", "agent_sha256": "sha256:0000000000000000000000000000000000000000000000000000000000000000", "benchmark_sha256": "sha256:0000000000000000000000000000000000000000000000000000000000000000", "harbor_runtime_sha256": "sha256:0000000000000000000000000000000000000000000000000000000000000000", "server_context_tokens": 65536, "max_output_tokens": 8192, "reasoning_required": true, "sample_task_count": 8, "sample_task_names": ["task-a", "task-b", "task-c", "task-d", "task-e", "task-f", "task-g", "task-h"], "sample_tasks_sha256": "sha256:0000000000000000000000000000000000000000000000000000000000000000" }, "objective": { "kind": "maximum_goodput", "maximum_error_rate": 0.0 }, "workload": { "kind": "benchmark", "sample_task_count": 8, "sample_task_names": ["task-a", "task-b", "task-c", "task-d", "task-e", "task-f", "task-g", "task-h"], "sample_tasks_sha256": "sha256:0000000000000000000000000000000000000000000000000000000000000000", "minimum_observations_per_point": 8, "boundary_repetitions": 3 }, "candidate_concurrency": [1, 2, 4, 8, 16, 32, 64], "points": [], "selection": null, "artifacts": { "bucket": "organization/benchmark-runs", "prefix": "serving-profiles/ornith-35b-q4-h200-64k-20260718" } } ``` The checked-in [`serving-profile-v1.schema.json`](../schemas/serving-profile-v1.schema.json) is the machine-readable contract. The [`serving-profile.json`](../examples/serving-profile.json) example validates against it. ## Identity And Reuse `model_sha256`, `deployment_sha256`, and `agent_sha256` are canonical profile digests. Together they cover the model repository and revision, weight format and quantization, serving engine and image, ordered arguments, hardware, replicas, activation and KV precision, context and batching limits, chat template, reasoning parser, sampling, caching, and speculative decoding. The deployment digest excludes the endpoint resource reference. A managed endpoint receives its deterministic name only after planning, and that transient address must not change the serving configuration identity. `harbor_runtime_sha256`, reasoning mode, exact sampled task names, and the sampled-task digest bind the selected concurrency to the exact Harbor client runtime and benchmark workload used to measure it. Plans may choose an explicit cohort when some benchmark tasks require environment capabilities unavailable on the profiling backend. That cohort is immutable evidence, not a runtime skip list. `benchmark_sha256` covers the benchmark revision, task digests, and the sampled workload distribution. It prevents a short synthetic sweep from being treated as proof of ShellBench task throughput. The profile's `plan.json` contains the resolved model, deployment, agent, and benchmark profiles used to derive these digests. The final profile remains small and queryable while the plan preserves every behavior-affecting value. A run may reuse a selection only when all four digests and both token limits match exactly. A runtime, quantization, template, reasoning, hardware, context, output, or benchmark change requires a new profile. A profile is not portable merely because the model name and GPU name are unchanged. ## Candidate Ladder Run a verified smoke request, then test concurrency in ascending powers of two: ```text c1, c2, c4, c8, c16, c32, c64, c128, ... ``` The ladder is not capped at `c64`. Extend it while aggregate throughput or goodput improves and the deployment remains stable. After identifying the last-good and first-bad powers of two, optional refinement points such as `c24` may be added between them. Use at least `max(8, 2 * concurrency)` observations at each point; the declared minimum is only a floor. Repeat boundary candidates at least three times. Keep client request concurrency, server sequence capacity, Harbor trial concurrency, active shards, and replica count separate in the evidence. Provider profiles use distinct benchmark tasks for every observation at a point so independent trials retain independent recorder retry budgets. Before the ladder starts, calibrated requests approach the declared context boundary while submitting the full declared output limit; a profile fails if either limit is not accepted. ## Capacity controls Profile these limits separately because they control different resources: - `trial_job_template.max_jobs` limits active trial Jobs for one Run; - namespace and hardware caps limit Jobs across Runs; - Job start pacing limits how quickly new Jobs are authorized; - `inference_max_concurrency` limits provider requests from one trial Job; - `inference_max_total_concurrency` limits provider request units reserved by one Run; and - budget admission limits work by reservation and cumulative ceiling. Compare trial Job limits at the same per-Job resources and provider limits. Later profiles may test higher Run, namespace, hardware, provider, or start-rate limits, but each tested point must name the limit that changed and keep the others fixed. Use ascending powers of two for each capacity boundary. Record successful terminal trials per hour as the primary goodput measure together with raw successes and attempts. Also record Job startup p50 and p95, queued time, provider throttling, infrastructure failures, cleanup failures, deadline headroom, active and reserved spend, and the limiting reason reported by the control service. Choose no production namespace, hardware, provider, or start-rate value from employee access or an undocumented assumption. The run plan must contain a verified quota or measured capacity source. It must also set a minimum worthwhile absolute goodput gain before the final comparison. A cleanup error, unsafe failure rate, provider throttling, missed deadline, or cost ceiling vetoes a candidate even when its throughput is higher. When candidates are within the practical tie range, select the lower limit. A profile can recommend a value but cannot promote the service capacity profile or change a run lock. Promotion and launch remain separate reviewed actions. ## Workload Profile against the workload the full run will run. For benchmark-speed selection, use a representative task sample with the same agent, tools, reasoning mode, context limit, output limit, sampling, and execution environment. When sampled tasks require a benchmark judge, profiling uses the exact judge configuration preserved in the run lock. The profiling recorder reads the locked provider's matching secret, forwards to the locked API URL and model, and enforces the locked reasoning-effort and temperature policy. The HF Job token remains the separate ingress credential used to reach the recorder. Profiling must not silently replace a direct OpenAI or Gemini judge with the Hugging Face router. Synthetic request tests may separately characterize prefill and decode capacity, but they do not select Harbor task concurrency by themselves. Record observed prompt and output token distributions rather than claiming that every request exercised the configured context limit. ## Measurements Each completed benchmark point records: - planned, completed, and failed request or task counts; - aggregate input and output tokens per second; - task completions per hour when profiling benchmark work; - per-session output tokens per second; - p50, p95, and p99 trial latency; - error and goodput rates; - peak device memory when observable; - the raw artifact prefix and checksum. It also records observed p50, p95, and maximum prompt and output tokens. These measure active workload shape and must not be replaced with configured limits. TTFT and TPOT are reported only when a streaming recorder measures them. A non-streaming full-response duration must never be labeled TTFT or TPOT. Successful HTTP responses are insufficient. The ladder runs the pinned Harbor task sample through the declared agent runtime, then verifies token accounting, task completion, endpoint logs, agent exits, truncation, timeouts, and hidden 4xx or 5xx responses. ## Stopping Rules Confirm and stop the ascending ladder when one of these conditions occurs: - allocation failure, OOM protection, or unsafe memory headroom; - repeated request failures, malformed output, timeouts, or endpoint errors; - aggregate throughput or goodput is flat or lower at two successive points; - declared p95 or p99 latency limits are exceeded; - per-session decode speed falls below the declared minimum; - queueing dominates without increasing completed work. Retry one failed point after a health probe before declaring the boundary. Keep failed and skipped points with explicit reasons. Never lower safety controls to force a larger concurrency result. ## Selection Choose one objective before the run: - `maximum_throughput`: greatest aggregate output throughput; - `maximum_goodput`: greatest completed work satisfying every declared limit; - `maximum_stable_concurrency`: highest repeatedly stable point; - `interactive`: greatest goodput satisfying interactive latency and per-session decode limits. The selection names the winning concurrency, criterion, supporting point digests, and rationale. The full run's `execution.concurrent_trials` must equal `selection.concurrency`. For capacity admission, the selection also records the tested worker, run Job, namespace, hardware, provider, and start-rate values. It records the effective concurrency, active limiting reason, raw completed and attempted counts, and minimum worthwhile effect. It distinguishes the measured recommendation from a later approved profile promotion. Higher concurrency is not automatically better. Prefer the lower point when two candidates are within measurement noise or the absolute gain is smaller than the registered minimum worthwhile effect. Safety, cleanup, provider, deadline, and cost vetoes take precedence over the primary metric. ## Storage Store profiles under one private Bucket prefix: ```text serving-profiles// plan.json points///evidence.json points///harbor-execution/ profile.json checksums.json _SELECTED | _FAILED ``` Write `profile.json`, `checksums.json`, and the terminal marker only after the endpoint is paused and reports zero ready replicas. The profile points and raw logs are immutable. A retry appends a new repetition or creates a new profile; it never overwrites prior evidence. The final run evidence records the profile Bucket URI and SHA-256 digest through `execution.serving_profile`. Manifest validation rejects a mismatched selection concurrency or serving identity before run planning. ## CLI The production CLI exposes: ```text harbor-hf profile plan EXPERIMENT --profile-id ID --max-spend-usd USD \ --estimated-profile-cost-usd USD \ --timeout-seconds 3600 --output plan.json harbor-hf profile preflight plan.json harbor-hf profile run plan.json harbor-hf profile select profile.json --output selected-profile.json ``` `plan` resolves exact identities and creates the candidate ladder without remote work. The plan embeds the immutable experiment, so the remote worker does not depend on mutable local state. `preflight` verifies the model revision, private Bucket, provider route or endpoint compute, current accelerator quota, hourly price, worst-case profile cost, and declared spend cap. Unknown endpoint quota fails closed. Provider profiles require an explicit estimate for the full profile through `--estimated-profile-cost-usd`; this is distinct from the deployment's run-wave estimate. Preflight rejects it when it exceeds either the provider or profile spend cap. Endpoint profiles omit this option. `run` submits one Hugging Face Job. For an Inference Endpoint, the worker requires a paused baseline, starts the cleanup watchdog before resume, keeps one endpoint lease across the whole ladder, and pauses and verifies zero ready replicas on every exit path. It first verifies ordinary chat, the reasoning channel when required, and a forced tool call. It then records content-free request observations for the endpoint or Inference Provider, tests ascending powers of two by running the sampled benchmark tasks through Harbor and the declared agent, and repeats the last two viable boundary points until each has three successful measurements. Failed health-check attempts remain in the raw evidence but do not replace those measurements. Provider points use the same distinct task set at every concurrency so workload composition cannot affect selection. Cleanup uncertainty leaves the profile nonterminal for operator or watchdog recovery. The operator machine never loads model weights or performs inference. The worker is restart-safe. It stores an absolute profile deadline before remote work, validates the original plan on restart, recomputes saved point metrics from their raw task observations, and resumes only missing ladder or boundary repetitions. If selection was already written, it verifies the selected profile against the plan and requires the endpoint to be paused before publishing `_SELECTED`. A conflicting plan, marker, point, or expired deadline still fails closed. This preserves the original cost bound across controller restarts instead of silently granting a fresh profiling budget. `select` recomputes every point digest before choosing the winner. A run can bind the resulting profile under `execution.serving_profile`; validation then requires exact model, deployment, agent, benchmark, context, output, and concurrency agreement. The binding is propagated into run locks and run digests. Do not split `profile run` into independent Jobs or runs per point. The profiler reuses one safely leased endpoint across the ladder and retains the watchdog and verified-pause guarantees. ## Not Covered This format does not define benchmark tasks, model quality, verifier scoring, autoscaling across multiple replicas, or provider pricing. It selects one serving configuration for one declared workload. The full benchmark remains the authority for task quality.