Plan the deployment and model portfolio
Deployment boundary planning maps hard constraints across tenants, regions, and networks before selecting a model pool shape and portfolio.
Plan deployment boundaries, pool shapes, route portfolios, endpoint independence, and installation handoff before approving model routing in Duale AI.
- Map hard boundaries for tenants, regions, credentials, networks, and data paths before comparing models.
- Choose pool shapes from dedicated tenants through multi-provider resilience to regional or on-premises deployments.
- Verify endpoint independence by recording shared and independent failure domains for each route.
- Prepare route records with USD pricing, capability evidence, credential ownership, and expiry conditions.
- Plan the installation handoff with topology, route inventory, network requirements, and version compatibility.
Summaries were generated by AI. Generative AI is experimental.
Plan the deployment boundary before you compare model benchmark scores. Provider choice is one part of a wider design that includes tenants, regions, credentials, networks, quotas, tools, data stores, monitoring, and support access.
Before approval, confirm that the workload is inside the documented scope and limits, and record any transparency or separate-agreement requirement. A human review step does not make an excluded use available.
Map non-negotiable boundaries
Write the hard boundaries down before you compare a model. A boundary discovered after approval can require a new tenant, provider account, credential, storage boundary, or deployment instead of a route edit.
For each workload, record:
- who needs isolated access or evidence, including customers and legal entities;
- where processing can occur, what cannot cross a border, and which data, retention, deletion, and human-access rules apply;
- which providers, accounts, contracts, and subprocessors are approved;
- where the data path must run, from collection and transit through compute, storage, cache, backup, and restore;
- where support, tools, and external recipients can access or send data;
- which output forms, context sizes, media types, tools, and document sources the workload needs;
- which availability, recovery, latency, and cost objectives apply; and
- who can approve a boundary change and who can stop the service.
Use separate deployments when the full processing path needs a different hosting, network, data, or operational boundary. A second platform instance can still share identity, telemetry, storage accounts, support services, or other landing-zone dependencies. A regional boundary needs region-local dependencies for the complete data path, not only another endpoint.
Use separate tenants for logical authorization, configuration, usage, cost views, conversations, and tenant-scoped Library authorization, records, and object namespaces inside one deployment. A tenant does not create a distinct provider account, credential, runtime, bucket, storage account, service identity, operator plane, or physical region. Provision a separate deployment, account, credential, or storage boundary when the requirement needs it.
Approve the boundary map only when every required path has an enforcement point and an owner. Then choose the pool shape that fits those boundaries.
Choose the pool shape
Start from the hard outcome, then choose the smallest pool that meets it:
- Required outcome
- Exact provider, model, or version
- Recommended shape
- Dedicated tenant with only standalone routes for the pinned identity
- Main trade-off
- Safe failure replaces fallback to another model
- Required outcome
- Multi-provider resilience
- Recommended shape
- Compatible routes with independent provider, account, quota, endpoint, network, and credential paths
- Main trade-off
- More choices do not guarantee a later attempt
- Required outcome
- Cost or latency flexibility
- Recommended shape
- Contract-compatible routes with truthful USD prices and representative performance evidence
- Main trade-off
- Routing remains contextual; preferences are not limits
- Required outcome
- Specialist work with a strict boundary
- Recommended shape
- Separate tenant or deployment for the specialist routes
- Main trade-off
- Skill preferences alone cannot exclude general routes
- Required outcome
- Customer-specific logical separation
- Recommended shape
- Separate tenant and, when needed, separate provider accounts and credentials
- Main trade-off
- A tenant does not give physical or regional isolation
- Required outcome
- Regional or on-premises boundary
- Recommended shape
- Separate qualified deployment with a complete local dependency and access path
- Main trade-off
- The full data path must meet the boundary, not only the model
The pool shape is valid only when every route it permits can serve every workload that reaches the pool. Split the pool when one route would violate a hard boundary or essential application contract.
Choose the portfolio
Approve a portfolio only when every enabled route meets the workload’s essential contract and every shared failure domain is understood.
Assess each route against the same evidence:
- Area
- Functional fit
- Questions to answer
- Does it support the response format, tools, input types, context, languages, and workload acceptance criteria?
- Area
- Data and contract
- Questions to answer
- Where does processing occur, what terms apply, and which data or retention options are allowed?
- Area
- Reliability
- Questions to answer
- Which provider, account, region, quota, gateway, network, and credential failure domains does it use?
- Area
- Performance
- Questions to answer
- What latency, throughput, rate limits, and context behavior did you measure on representative work?
- Area
- Cost
- Questions to answer
- What are the input, output, cached-input, retry, fallback, and supporting-service costs?
- Area
- Lifecycle
- Questions to answer
- How are versions named, deprecated, upgraded, disabled, and supported?
- Area
- Assurance
- Questions to answer
- Which evaluation, security, privacy, legal, and operational reviews support approval?
Provider capability and compatibility states what each provider type accepts and how you declare it. Understand model routing states what a wrong declaration costs.
A resilient pool can contain two or more independent compatible routes and optional specialist or lower-cost routes. Duale AI ranks eligible routes for each request. Route names and configuration order do not define primary or fallback roles. Whether and when a later attempt runs depends on current routing conditions, the deadline, and available attempt capacity. It can start before another attempt fails.
For an exact-model boundary, include only routes whose provider identity and version contract meet the requirement. You can include several endpoint routes for that same model version when the account, region, and credential boundaries allow them. Keep one route only when those boundaries also require one exact endpoint.
Verify endpoint independence
Configure each independently selectable endpoint as a separate route with a stable, unique model_name inside the
tenant. A shared provider account, regional quota, gateway, network path, credential, or upstream host removes
independence for that failure, but the routes can still cover other failures. Record the exact shared boundary instead of
calling the routes fully independent.
Routes that use the same provider type and endpoint can also share endpoint health state. A different route name or model ID, tenant, or credential does not make that health independent within one deployment.
For each route, record the shared and independent failure domains. Test the terminal result when the first-ranked route fails, a later route succeeds, and a blocking error or deadline prevents another attempt. Retries can add calls, quota use, cost, and delay. Multiple routes provide choices, not a fixed traffic share or availability level.
Prepare each route record
Check the tenant-wide currency constraint before you complete the route record:
All enabled routes currently require pricing in USD. The schema accepts another currency, but one such route stops routing for every request in that tenant until you correct or disable it. Check the currency before you enable the first route.
Before approval, record:
- tenant, deployment, provider account, endpoint, region, and stable
model_name; - provider legal entity, model identifier, version contract, and routing or inference profile;
- supported functions, context and output limits, prices, quotas, and capability evidence;
- provider terms, retention, training, subprocessors, support locations, and incident process;
- credential owner, rotation method, network path, and emergency cutoff;
- change owner, evidence date, review or expiry condition, and replacement plan.
Unknown evidence must stay unknown. Do not enter a free price, supported capability, or successful benchmark merely to complete a field. Store governance fields in the customer change system and link them to the configuration identifier; they are not model-route configuration fields.
A complete route record tells the reviewer what is approved, when its evidence expires, and which control stops the route.
Create a route
A saved route is a configuration record, not evidence that the provider works. Create it in this order:
- Choose the deployment and tenant. Open the Dashboard connected to the target deployment, then select the tenant. The tenant selector does not switch deployments.
- Protect configuration access. Restrict
config:read_config,config:create_config,config:update_config, andconfig:delete_configto authorized users. A read returns the complete route document, including provider credential fields, so treat it as secret-bearing access. - Start with the route disabled. Open Settings → LLM Providers, then select Create or Create from preset. A route defaults to enabled when the field is absent. Turn Enabled off before the first standalone save. For a preset, select Detach to edit for Enabled, turn it off, then save.
- Enter the route identity and evidence. Use
model_nameas the stable unique route name,display_nameas the human label, andmodel_idas the provider-facing identifier. Add the approved provider identity, performance limits, USD pricing, benchmarks, capabilities, and every required secret or approved ambient identity. Review every effective field a preset supplies. - Supply and test the connection path. Most provider types use
base_url; AWS Bedrock and Vertex AI have their own location and endpoint fields. A saved endpoint does not create DNS, TLS, proxy, firewall, VPN, or private-link reachability. The deployment and network owners must supply and test that path. - Verify the resolved value and locality evidence. Reject omitted, ambiguous, global, or cross-region profiles when locality is a hard boundary. A route region and a valid form do not prove the processing destination.
Provider capability and compatibility lists the supported provider types and the fields that affect each one.
Saving validates the form’s shape and nothing else. It does not contact, prewarm, or qualify the provider, and missing evidence does not quarantine an enabled route. Complete live technical and business evaluation before production use.
Plan the installation handoff
Customer-hosted deployment is available to eligible customers. Write to contact+security@mail.duale.ai to confirm eligibility and receive the supported deployment specification. That specification defines the platform release, topology, version compatibility, connectivity, migration, upgrade process, support, and the responsibility split for the installation. Assemble the customer-owned requirements below before the deployment review.
Include the following information in the architecture handoff:
- the supported topology, deployment build, tenant and agent provisioning sequence, and tenant map;
- the approved route inventory, expected fallback behavior, and shared failure-domain assumptions that must be tested;
- network, private-connectivity, identity, credential, storage, capacity, quota, telemetry, backup, support, and incident requirements;
- complete task API base URL with any gateway prefix and one-time token delivery;
- tenant and agent identifiers, grants, local-cache paths, telemetry settings,
same-origin
/librariesaccess, and every presigned storage host that the client must reach; - qualified SDK and platform versions by function and their upgrade order;
- the change, evidence, exception, and retirement process.
Do not approve the installation until the supported specification maps every handoff item to an owner, interface, and testable result.
Verify the architecture
The architecture is ready only when each hard boundary has an owner, an enforcement point, and observed evidence.
The plan is ready for approval only when:
- every hard requirement maps to a deployment, tenant, route, provider, network, application, or operating control;
- every potentially selectable route is allowed for every task that can reach its tenant pool;
- the application has passed live outcome and failure tests against each required route combination;
- the failure-domain review shows which dependencies remain shared;
- the complete data-flow and human-access evidence supports each locality or customer claim;
- each deployment bundle, owner, change gate, rollback, and emergency cutoff is named; and
- the resolved configuration is treated as desired state, not proof that every runtime instance has applied it.
If any item lacks evidence, keep the plan unapproved and record the missing owner or contract. When the checks pass, Ownership and approvals defines the decision record.