PRICING

    Pricing built around how you run inference.

    Start instantly with usage-based Endpoints, deploy dedicated Sovereign Endpoints™, or run the InferX Platform inside your own infrastructure.

    WAYS TO GET STARTED

    Choose the path that fits your build.

    Start small, build regularly, or design a production deployment with the InferX team. Pay-as-you-go continues after included credits are used; it is not a separate plan.

    FREE

    Evaluate InferX.

    Try production endpoints and APIs with a small amount of usage included.

    $10in free usage credits
    • $10 usage credits included
    • Try hosted production endpoints
    • OpenAI-compatible API access
    • No commitment

    After free credits are used, continue with pay-as-you-go pricing.

    Start Free
    ENTERPRISE

    Design for production.

    Need dedicated infrastructure, enterprise support, or custom deployments? We'll help design the right InferX deployment for your organization.

    Custom Pricingfor production AI deployments
    • Dedicated deployments
    • Enterprise SLAs
    • Private networking
    • Volume pricing
    • Priority support
    • Architecture guidance
    Contact Sales

    Free and Developer usage credits do not represent a separate pay-as-you-go plan. Billing continues at published usage rates once the applicable included credits are exhausted.

    INTRODUCTORY MODEL PRICING

    Introductory pricing compared with common market API rates.

    Representative hosted endpoint rates for popular ready-now models.
    ModelMarket inputInferX inputCached inputInput discountMarket outputInferX outputOutput discount
    DeepSeek V4 Flash$0.14 / 1M$0.07 / 1M$0.01 / 1M50%$0.28 / 1M$0.11 / 1M61%
    Mimo v2.5$0.14 / 1M$0.08 / 1M$0.008 / 1M43%$0.28 / 1M$0.18 / 1M36%
    GLM 5.2$1.40 / 1M$0.18 / 1M$0.018 / 1M87%$4.40 / 1M$1.90 / 1M57%

    InferX supports cached-input pricing for repeated context and cache-hit workloads.

    SERVERLESS ECONOMICS

    Serverless compute vs always-on GPU hosting.

    Pay for active inference compute instead of idle capacity.
    THE OLD WAY

    Dedicated H100

    Always on. Always billing.

    Hourly rate
    $4.00 / hr
    Hours / month
    730 hrs
    Calculation
    $4.00 × 730
    $2,900per month

    Paying for idle capacity.

    THE INFERX WAY

    InferX Serverless

    Active compute only. Scale to zero between requests.

    Active GPU time
    $3.95 / GPU hr
    Standby cost
    Scale to zero
    Reference workload
    Bursty usage
    ~$220per month

    Pay only for active compute.

    ~13×lower reference monthly cost$2,900 → ~$220for the same production workload pattern
    Illustrative comparison based on the original InferX economics diagram. Actual cost varies by model, GPU class, active compute time, and deployment configuration.
    SERVERLESS COMPUTE

    Serverless inference compute pricing

    For custom deployments and dedicated inference endpoints, InferX bills active GPU compute time by the second. You pay for the compute your inference workload actually uses, not idle capacity.

    H100$3.95 / GPU-hour about $0.00110 / GPU-second
    H200$3.95 / GPU-hour about $0.00110 / GPU-second
    B300$6.95 / GPU-hour about $0.00193 / GPU-second
    This is for custom deployments / dedicated inference endpoints. Hosted pay-per-token API pricing is separate.

    MONTHLY ACCESS FAQ

    Simple access, clear billing.

    Are credits included every month?

    Yes. The $50 in usage credits are included with each monthly billing cycle.

    What happens after I use the credits?

    You can continue using InferX seamlessly with pay-per-usage billing.

    Do unused credits roll over?

    Unused credits do not currently roll over.

    Can I cancel anytime?

    Yes. You can cancel anytime.

    Which InferX products can credits be used for?

    Credits can be used for any InferX product.