Tiers #
Prices are per month in Canadian dollars, billed monthly from the day a tier is activated. The shared tier is included with preview access. GPU capacity is limited: a GPU tier request can wait before it is provisioned, is not billed while it waits, and the administrator tells you when the GPUs are available.
Shared
Included with preview access
The preview default: every organisation shares the Falcon API and the model services, metered by the quota buckets.
- Every public route, metered per quota bucket
- The Falcon API and the Express Voice service kept warm; Express Cue, the Waymark service (Waymark and Waymark Extra) and Cognitio scale to zero after a quiet period, so the first call after one waits for the service to start
- No commitment; included with preview access
Tier id shared
Dedicated Warm
249 CAD a month
A private always-warm Falcon API and Express Cue instance for one organisation, with no scale-to-zero wait. The in-process models run privately when the organisation calls its private API address.
- A private Falcon API instance kept warm 24/7 at its own address, serving the in-process models (ELIM, LIM, LIM Nano, LIM3D, LIM3D-XL) to the requests sent there
- A private Express Cue instance kept warm, used automatically for the organisation’s keys
- No scale-to-zero wait on the private instances
- The shared GPU services for the routes not hosted privately
Tier id dedicated-cpu
Dedicated GPU
1,499 CAD a month
One private NVIDIA L4 kept warm 24/7, hosting one GPU model service of the organisation’s choice (Express Voice, Waymark with Waymark Extra, or Cognitio), plus everything in Dedicated Warm.
- Everything in Dedicated Warm
- One NVIDIA L4 kept warm 24/7, private to the organisation
- One GPU model service on the L4: Express Voice, Waymark (both Waymark models) or Cognitio
- No scale-to-zero wait on the hosted GPU model
Tier id dedicated-gpu
Dedicated GPU ×2
2,799 CAD a month
Two private NVIDIA L4s kept warm 24/7 for organisations that host two GPU model services, one on each L4, plus everything in Dedicated Warm.
- Everything in Dedicated Warm
- Two NVIDIA L4s kept warm 24/7, private to the organisation
- Up to two GPU model services, one per L4, from Express Voice, Waymark (both Waymark models) and Cognitio; a single chosen service runs on both L4s
- No scale-to-zero wait on the hosted GPU models
Tier id dedicated-gpu-2
What runs where #
Where each service answers on each tier.
| Service | Shared | Dedicated Warm | Dedicated GPU | Dedicated GPU ×2 |
|---|---|---|---|---|
| Falcon API with ELIM, LIM, LIM Nano, LIM3D, LIM3D-XL, called at the public base URL | Shared | Shared | Shared | Shared |
| The same models, called at the organisation’s private API address | None | Private, kept warm | Private, kept warm | Private, kept warm |
| Express Cue | Shared | Private, kept warm | Private, kept warm | Private, kept warm |
| GPU model services: Express Voice, Cognitio, and Waymark with Waymark Extra as one service | Shared | Shared | One chosen service private; the rest shared | Up to two chosen services private; the rest shared |
| Private NVIDIA L4 GPUs | None | None | One | Two |
Express Cue and the GPU models a tier hosts answer from the organisation’s own services whichever address its keys call.
Each L4 hosts one GPU model service: Express Voice, Waymark or Cognitio. Waymark and Waymark Extra run together as one service, as they do on the shared fleet, so choosing Waymark hosts both. Dedicated GPU hosts one of the three. Dedicated GPU ×2 hosts up to two, one on each L4; if only one is chosen, it runs on both L4s for more throughput. On the shared services the Falcon API and Express Voice are kept warm; Express Cue scales to zero after a quiet period and takes a moment to start again, and the Waymark service and Cognitio scale to zero too, so their first call can take one to two minutes.
What changes once a tier is active #
- Calls made with the organisation’s keys to Express Cue and to the GPU models the tier hosts go to its own services, at the public base URL as at the private one. The public routes, request and response shapes, error codes and body limits are unchanged.
- The Falcon API’s in-process models (ELIM, LIM, LIM Nano, LIM3D and LIM3D-XL) run privately for requests sent to the private API address. To use it, change the base URL of the integration to that address; the keys stay the same. Requests to the public base URL keep running those models on the shared Falcon API.
- Calls are still counted per quota bucket so that the dashboard shows usage. The sandbox works as before through the public base URL, so its Express Cue and hosted GPU model calls reach the dedicated services.
- A GPU model the organisation did not choose for its tier, and every route on the shared tier, keeps going to the shared services with their usual behaviour.
- The dashboard shows the tier, the date it was activated, the models hosted and the endpoints in use, the private API address among them.
Provisioning and lead time #
- Sign in to the portal and open the dashboard. The Hosting panel shows the current tier and its status.
- Choose a tier in Request dedicated hosting. A GPU tier asks which models to host; add notes — expected traffic, peak hours, who to contact — and send the request.
- The panel shows Requested with the date. A Falcon administrator reviews the request; until the panel shows Active you can cancel it from the same panel and stay on the shared tier.
- Once approved, the administrator provisions Dedicated Warm within one business day, and a GPU tier when GPU capacity is available; the panel then shows Active with the activated date and the endpoints. Express Cue and the hosted GPU models switch over with nothing to change on your side.
Requests are approved by an administrator, not self-served; the Portal page describes what the administrator sees. An organisation with a dedicated tier already active cannot send a second request (409 hosting_active): ask an administrator to change the tier instead.
Billing and cancellation #
A dedicated tier is billed monthly in CAD from the day it is activated. Cancel at any time for the end of the current month by asking an administrator; the organisation then returns to the shared tier and its keys route to the shared services again, with no change to the keys themselves. The private API address stops answering, so an integration pointed at it moves back to the public base URL. Moving between dedicated tiers takes effect when the administrator saves the new tier and is billed from the next month at the new price.
Limits #
- GPU capacity is subject to availability. A GPU request can wait for hardware; the panel stays at Requested until it is provisioned.
- One region: every dedicated instance runs in
us-central1, the region of the shared fleet. - Each L4 hosts one GPU model service (Express Voice, Waymark with Waymark Extra, or Cognitio): one on Dedicated GPU, up to two on Dedicated GPU ×2, where a single choice runs on both L4s. The administrator confirms the choice before provisioning.
- The in-process models are private only at the private API address; the public base URL keeps serving them from the shared Falcon API.
- The models are the ones on this site at their published versions; Waymark Flight is not offered on dedicated tiers and stays on the shared fleet. Fixed weights per version; the model does not learn from requests, on a dedicated tier as on the shared one.
- The private preview terms apply: weights are not distributed, and a dedicated instance is operated by Ducky Software, not handed over.