Private LLM deployment in your own cloud account

Deploy your own private LLM, on infrastructure you own.

LLM Forge provisions a GPU server in your AWS or RunPod account and gives you a private, key-authenticated OpenAI-compatible endpoint. Your prompts and responses stay on your instance. Delete it when you are done and we verify nothing is left running.

Start 24-hour trial Read the docs
Card required · full Lab features · 24 hours
BEFORE YOU SPEND ANYTHING
The exact instance, region and hourly rate
A confirmation page lists every resource we will create in your account, before we create it.
WHILE IT RUNS
A hard budget cap and a self-destruct timer
Required on every deployment. You can raise the cap; you cannot disable it. At 100% we destroy the deployment.
WHEN YOU'RE DONE
Delete everything, verified clean
After teardown we sweep the region and show the result: zero resources remaining.

A full LLM stack, ready in minutes

typically 5 to 25 minutes from click to first token
Endpoint gateway
Key-authenticated HTTPS, OpenAI-compatible routes, streaming responses.
/v1/chat/completions
Inference engine
Batching, concurrency limits and configurable context length, tuned per model.
8,192 ctx default
Model weights
Ten curated open models, downloaded onto your instance disk. Optional knowledge files.
Qwen3 · Llama · Mistral
GPU instance
Single or multi-GPU, the tier and region you picked, with a budget cap and optional timer.
L4 · L40S · multi-GPU
Your cloud account
Every resource tagged and owned by you, created through a role you can revoke.
AWS · RunPod
WHAT YOU DO
01Connect a cloud account, once.
02Pick a model and a hardware tier.
03Set a budget cap and confirm the estimate.
04Copy the endpoint URL and key into your code.
No Terraform to write, no AMI to build, no engine to configure. Nothing to uninstall afterwards.

Own your infrastructure

The instance, the disk, the network and the endpoint belong to your account. You can open them in your own console, put them behind your own policies, and keep them if you stop using us.
No lock-in: resources stay after you cancel
Your provider rates, your reserved capacity, your credits
Access granted by a role you can revoke in one command

Monitor and log end to end

Full visibility of what, where, when and how. Every action we take in your account is written to a timestamped log, including the plan of what will be created before it exists.
WHEN
WHAT
WHERE
08:03:08
Plan recorded: 1 instance, 1 security group
us-east-1
08:04:20
GPU instance created, billing starts
g6e.xlarge
08:19:12
Endpoint live and answering
203.0.113.10
21:44:03
Destroyed and swept, zero resources remaining
us-east-1
Exportable as JSON, kept after teardown, and paired with spend against cap and request counts read from your own gateway.

Private by construction

not a setting you have to find
Your model, your machine
Weights are downloaded onto a GPU instance inside your own cloud account. No shared tenancy, no queue behind other customers, no third party holding your model.
Prompts never reach us
Requests go straight from your client to your endpoint. We do not proxy, store, log or train on your content, and we could not read it if we wanted to.
EU-only infrastructure
EU
One checkbox pins every resource to EU member-state regions and keeps it there.
eu-central-1 eu-west-1 eu-west-3 eu-north-1 eu-south-1 eu-south-2
Access you control
One endpoint key, shown once and stored only as a hash. Rotate it or delete the deployment and it stops working immediately. On AWS we borrow a role for an hour at a time, scoped to resources we tagged; delete the stack and our access ends.
BUILT TO MEET
GDPR data residency No sub-processing of prompt data Encryption in transit Full audit trail Verified deletion
Your prompt and response data never enters our systems, so it is never processed, stored or shared by us.

Pricing

One flat subscription for the platform. GPU hours are billed to you by your cloud provider.

Lab
For evals, prototypes and side projects.
$49/month
Start 24-hour trial
2 concurrent deployments
AWS and RunPod
Single-GPU hardware, up to 48 GB
30 days of history after teardown
Email notifications and support
Platform
COMING SOON
Deploy your own LLM SaaS.
$299/month
Join the waitlist
EVERYTHING IN PRO, PLUS
Your customers' domains Share links Per-customer keys Teams and roles
Tell us what you need and we will factor it into the build.

Two bills, two companies

Your plan covers the platform: orchestration, dashboard, budget caps and teardown. The GPUs run in your own cloud account and your provider invoices you directly for them. We do not mark up, resell or take a cut of compute, and we cannot refund it.
FROM US
$99/mo, flat
The same whether you deploy once a month or every day.
FROM YOUR CLOUD
$1.861/hr est.
Charged only while a GPU is running. Example: an L40S running Qwen3 8B.
COMPARE PLANS
Lab
ProFEATURED
Platformcoming soon
Price
$49/mo
$99/mo
$299/mo
Concurrent deployments
2
6
TBC
Cloud providers
AWS + RunPod
AWS + RunPod
TBC
Hardware
Single-GPU, up to 48 GB
+ multi-GPU, 70B-class
TBC
History after teardown
30 days
Full
TBC
Notifications
Email
Email + webhooks
TBC
Support
Email
Priority
TBC
On the way
None
None
Your customers' domains · share links · per-customer keys · teams
Included on every plan, including the trial
not upsells or add-ons
Hard budget caps
Required on every deployment. You can raise the cap; you cannot disable it.
Self-destruct timers
Set a lifetime up front so a forgotten GPU cannot run all weekend.
Verified teardown
A sweep after every destroy, with a record of what was removed.
Full audit trail
Every action we took in your account, timestamped, including the pre-flight plan.
EU-only residencyEU
One checkbox pins a deployment to EU-member-state regions.
Never-gated delete
Delete any deployment, or your whole organization, at any time on any plan.
Privacy and safety are not premium features.
How does the 24-hour trial work?
The trial runs for 24 hours. A card is required to start it and the plan you picked begins when the trial ends. Nothing is destroyed at the end of the trial: anything still running keeps running and keeps billing to your cloud account.
Do you resell GPU capacity?
No. Everything runs in your account at your provider's rates. We never see your infrastructure invoice.
What access do you need to my cloud?
On AWS, a role we can borrow for an hour at a time, scoped to resources we tagged ourselves. Delete the CloudFormation stack and our access is gone instantly.
What if a deploy fails halfway?
We explain the cause in plain language, remove whatever was created, and show the sweep result. The most common first failure is an AWS GPU quota of zero, which comes with a direct link to the form that fixes it.

Your first private model runs in about nine minutes.

24 hours of Lab. Card required. Connect a cloud account and pick a model.