Configure backend hosts and optional named path routes using semicolon-separated connection strings.
TL;DR
- Existing
Host1…HostNconfigurations require no changes — omitted mode retains probed/APIM behavior, accepts every priority in group1, and calls the configured host itself.- Use named
Host_*andPath_*settings for reusable routes — the longest matching prefix owns the request.- Host mode is explicit —
apimis a probed gateway,directis called without probes, andindirectis delegated throughvia.
Units: all timeouts in milliseconds unless noted. Delimiters:
;or,(both accepted).
| Key | Default | Description |
|---|---|---|
host |
(required) | Backend base URL. Protocol defaults to https:// if omitted. Trailing slashes are stripped. |
probe |
echo/resource?param1=sample |
Health probe path. Used only by mode=apim; ignored by direct and indirect. |
path |
/ |
Path prefix used for routing. Requests matching this prefix are sent to this host. |
acceptablePriorities |
* |
Colon-separated numeric request priorities accepted by this host, for example 1:2. * or omission accepts all priorities. |
priorityGroup |
1 |
Positive integer failover group. Lower groups are exhausted before higher groups. LoadBalanceMode orders peers within one group. |
via |
(empty) | Named mode=apim gateway host such as Host_apim. REQUIRED with mode=indirect and invalid with other modes. |
mode |
apim |
apim calls and probes this host; direct calls it without probes; indirect never calls or probes it and delegates through via. Unknown values are rejected. |
ipaddress |
(empty) | Override DNS — force all requests to this IP. |
processor |
(empty) | Custom stream processor name. Required and auto-defaulted in direct mode. |
usemi / useoauth |
false |
Attach a Managed Identity / OAuth2 Bearer token to every request and probe. |
audience |
(empty) | OAuth token audience. Required when usemi=true. |
authprovider |
AzureProvider |
Exact simple class name of an IBackendTokenProvider registered by AuthProviders. Matching is case-insensitive; no class-name suffix is implied. |
api-key |
(empty) | API key value to send on every forwarded request and probe. Sets auth mode to API key. |
api-key-header |
api-key |
Header name used when api-key is set. |
stripprefix / strippathprefix |
true |
Strip the matched path prefix before forwarding. Set false to preserve the full original path. |
retryafter / useretryafter |
true |
Honour the Retry-After header returned by the backend. |
Warning
An unrecognised or invalid key rejects the complete candidate host/route snapshot. A warm refresh retains the last known-good snapshot; an invalid initial configuration starts with no active hosts.
Rule: Use the connection string format for all new hosts — it keeps every option for a host in one variable.
host is the destination's base URL. Set ipaddress only when requests must reach a specific address without using DNS resolution for that host, for example when testing a selected backend instance. Leave it empty to follow DNS changes; a pinned address does not update when the DNS record changes.
Tip
A host fails after an address change? Check for an old ipaddress override. Remove it to resume DNS resolution, or update it to the intended address.
# Minimal — standard probed host
Host1="host=https://api.backend.com;probe=/health"
# Path-routed host (strip prefix, default)
Host2="host=https://chat-service.internal;path=/chat;probe=/health"
# Preserve full path (backend owns its own routing)
Host3="host=https://passthrough.internal;path=/api/v1;stripprefix=false"
# Authenticated host (Managed Identity)
Host4="host=https://secure-api.internal;usemi=true;audience=api://my-app-id;authprovider=AzureProvider;probe=/health"
# Authenticated host (API key with custom header)
Host5="host=https://secure-api.internal;api-key-header=foo;api-key=bar;probe=/health"
# Direct mode — serverless, no probing
Host6="host=https://my-func.azurewebsites.net;mode=direct;path=/api/v1"
# IP override — skip DNS
Host7="host=https://api.backend.com;ipaddress=10.0.1.5;probe=/health"Choose authentication per host so the proxy can meet each backend's requirements without callers supplying backend credentials. Managed identity avoids storing an API key; API-key mode supports backends that require a shared secret. Auth is configured through the host connection string, not a global UseOAuth switch.
host identifies where to send the request. audience identifies the resource the token is intended for; it is not necessarily the backend's hostname. With usemi=true, set the audience required by that backend and select a registered authprovider to acquire the token.
host=https://example.openai.azure.com
usemi=true;audience=https://cognitiveservices.azure.com
authprovider=AzureProvider
These lines describe fields in one host connection string. For Azure OpenAI, the token resource is https://cognitiveservices.azure.com even though the endpoint has its own hostname.
Tip
Backend returns 401 or 403? Check the expected token audience and the identity's backend permissions. Changing the destination hostname does not automatically change the audience or grant access.
| Host connection string values | Effective auth mode |
|---|---|
useoauth=true (or usemi=true) |
OAuth2 / Managed Identity |
api-key=<non-empty> |
API key mode (<api-key-header>: <api-key>) |
useoauth=false and empty api-key |
No auth header added |
For OAuth2 mode, authprovider selects one of the implementations registered globally through AuthProviders. The value MUST equal the implementation's simple runtime class name, such as AzureProvider or ContosoTokenSource; custom class names do not need to end in Provider.
Example custom header mapping:
Host1="host=https://example.internal;api-key-header=foo;api-key=bar"This sends foo: bar to that backend.
Note
Set only one auth mode per host. If both OAuth and API key entries are present, the host string order can change which mode is applied.
Note
Legacy format (Host1=https://..., Probe_path1=/health, IP1=10.0.1.5) is still supported but cannot express path, mode, usemi, or other per-host options. Do not mix legacy and connection-string keys for the same host number.
Rule: Set the mode according to who receives the proxy's HTTP request.
Host_apim="host=https://gateway.azure-api.net;mode=apim;probe=/status"
Host_direct="host=https://model.example.net;mode=direct"
Host_logical="host=https://ptu.example.net;mode=indirect;via=Host_apim"| Mode | Called by proxy? | Probed by proxy? | Required companion setting |
|---|---|---|---|
apim |
Yes | Yes | A usable probe path |
direct |
Yes | No | None |
indirect |
No | No | via=Host_<gateway> targeting an apim host |
Omitting mode preserves the existing apim behavior. A via target must use mode=apim; indirect-to-indirect chains, via on direct/APIM hosts, and indirect hosts outside a Path_* route are rejected.
Tip
Troubleshooting: If a candidate snapshot is rejected after adding via, verify that the logical host uses mode=indirect and the referenced gateway uses mode=apim.
Rule: Use mode=direct for any backend that scales to zero — the proxy will never probe it, so it will never wake it unnecessarily.
Host6="host=https://my-func.azurewebsites.net;mode=direct;path=/api/v1"In direct mode:
- No health probe is ever sent.
- The host is always treated as healthy (
SuccessRate = 1.0). - Average latency defaults to
0, so direct-mode hosts sort first inlatencyload-balance mode. processoris auto-set to the default stream processor if not specified.
Tip
Troubleshooting: If a direct-mode host starts returning errors, the circuit breaker still tracks failures per request — the host will be excluded once it breaches CBErrorThreshold.
Rule: Specific-path hosts always win over catch-all hosts; within matched hosts the load balancer decides.
Host1="host=https://chat-service.internal;path=/chat"
Host2="host=https://embed-service.internal;path=/embeddings"
Host3="host=https://default-service.internal" # catch-all (path=/)| Incoming request | Matched host | Forwarded path (stripprefix=true) |
|---|---|---|
GET /chat/completions |
Host1 | GET /completions |
POST /embeddings/create |
Host2 | POST /create |
GET /models |
Host3 | GET /models |
Path matching rules:
- Hosts with an explicit
pathprefix are checked first. /,/*, or emptypathis a catch-all and is tried only when no specific path matches.- Wildcards (
/api/*) match the same as the bare prefix (/api).
Note
stripprefix=false preserves the full original request path on the forwarded request. Use this when the backend application handles its own sub-routing under the same prefix.
Rule: A Path_* setting owns its longest matching prefix and references named hosts without duplicating their connection settings.
Host_chat_east="host=https://chat-east.internal;mode=direct;acceptablePriorities=1:2;priorityGroup=1"
Host_chat_fallback="host=https://chat-fallback.internal;mode=direct;acceptablePriorities=1:2:3;priorityGroup=2"
Path_chat="prefix=/api/chat;hosts=Host_chat_east:Host_chat_fallback;stripprefix=true"Named route fields:
| Field | Default | Description |
|---|---|---|
prefix |
(required) | Segment-boundary request prefix. /api matches /api and /api/x, not /apix. |
hosts |
(required) | Colon-separated Host_*, Host-*, or HostN references. |
stripprefix / strippathprefix |
true |
Remove the matched prefix before forwarding. |
The longest prefix wins. If that route has no host accepting the request priority, the proxy returns no candidates and does not fall through to a broader route. Hosts referenced by named routes are not also exposed through legacy catch-all selection.
Tip
Troubleshooting: If a route keeps the previous configuration after refresh, check for a missing host reference, duplicate prefix, mixed direct and via hosts, or an invalid via target. The rejected update is logged and the active snapshot remains unchanged.
Rule: Filter by acceptablePriorities, exhaust the lowest eligible priorityGroup, then advance to the next group.
Host_ptu="host=https://ptu.internal;mode=direct;acceptablePriorities=1;priorityGroup=1"
Host_paygo="host=https://paygo.internal;mode=direct;acceptablePriorities=1:2:3;priorityGroup=2"
Path_models="prefix=/models;hosts=Host_ptu:Host_paygo;stripprefix=false"For priority 1, PTU is tried before PayGo. Priorities 2 and 3 start directly at PayGo because no group-1 host accepts them. latency and timetofirstbyte sort only within a group, so a faster group-2 host cannot precede group 1.
Tip
Troubleshooting: A 503 with no attempted backend means the matched route has no host whose acceptablePriorities contains the resolved numeric request priority.
Rule: Set every logical APIM-selected backend to mode=indirect;via=Host_<gateway>; keep via off Path_*.
Host_apim="host=https://gateway.azure-api.net;mode=apim;probe=/status"
Host_ptu="host=https://ptu.openai.azure.com/openai;mode=indirect;via=Host_apim;acceptablePriorities=1;priorityGroup=1"
Path_openai="prefix=/api;hosts=Host_ptu;stripprefix=true"All hosts in one route must either be callable hosts or be indirect hosts referencing the same apim gateway. Missing, self-referencing, chained, mixed-mode, and multiple-gateway configurations are rejected atomically. The gateway owns proxy health and circuit state; indirect logical backends are not probed or called directly.
Note
In the current migration phase, via selects the APIM transport but APIM still reads its endpoint catalog and retry rules from the deployed policy fragment. The signed proxy-to-APIM route envelope and policy parser remain deferred work.
Tip
Troubleshooting: If APIM selects an unexpected endpoint, compare the logical host settings with the deployed APIM fragment. Until envelope support is completed, the fragment remains authoritative inside APIM.
Rule: The poller runs every PollInterval ms; a host is active only while its rolling success rate is ≥ SuccessRate%.
Every PollInterval ms:
For each probed host:
GET <ProbeUrl> (timeout = PollTimeout ms)
├── 2xx → AddCallSuccess(true) → latency recorded
└── else → AddCallSuccess(false) → latency not recorded
FilterActiveHosts:
active = hosts where SuccessRate() >= threshold
if latency order changed → invalidate shared iterator cache
| Config | Default | Description |
|---|---|---|
PollInterval |
15000 ms |
How often each host is probed |
PollTimeout |
3000 ms |
Max wait for a probe response |
SuccessRate |
80 % |
Minimum rolling success rate to stay active |
Note
Direct-mode hosts skip GetHostStatus entirely — they always return true and are included in FilterActiveHosts unconditionally.
Tip
Troubleshooting: If all hosts fall below the threshold the proxy returns 503. Lower SuccessRate or increase PollTimeout if backends are slow but functional.
Setup: 3 hosts,
LoadBalanceMode=latency,SuccessRate=80,PollInterval=15000.
| Host | Probe result | Rolling rate | Active? | Avg latency |
|---|---|---|---|---|
chat-service |
9/10 success | 90% | Yes | 120 ms |
embed-service |
6/10 success | 60% | No | — |
func-direct |
mode=direct |
always 100% | Yes | 0 ms |
In latency mode, func-direct (0 ms) is tried first, then chat-service (120 ms). embed-service is excluded until its rolling rate recovers above 80%.
- LOAD_BALANCING.md — How hosts are ordered and retried per request
- CIRCUIT_BREAKER.md — Per-request failure tracking and circuit state
- CONFIGURATION_SETTINGS.md —
PollInterval,PollTimeout,SuccessRateconfig keys - Customize a Backend Token Provider — Implement, register, and select an outbound token provider