
The developer experience and API are good enough. I mostly work with node and typescript for my backend and their official SDK makes it really easy to integrate so no need to worry about provisioning ec2 instances fighting with Cuda drivers or containerizing models. I can just grab llama-3 or Krea drop in my API keys and call one async function. The dashboard is super clean too, every model has this run with API tab that spits out ready to copy snippets with the schema already mapped. It basically deletes the Devops side of AI models so one can just focus on app logics. Review collected by and hosted on G2.com.
The cold boot time are definitely the biggest issue. If you use a slightly obscure model or custom fine tune that has not been hit in the last few minutes, the instance scale to zero. When the next request hit, it can easily take 10 to 30 seconds for the container to spin back up. Also the per-second compute pricing is fine for prototyping but scales up crazy fast if you get a traffic spike so you have to watch your billing limits. Review collected by and hosted on G2.com.