We're thrilled to launch Serverless model serving, offering elastic scaling, automatic load balancing, and a pay-as-you-go billing for model inference. With one-click deployment, users can seamlessly deploy models using public or private images, ensuring high availability and efficient performance at any scale. Billing is based on actual pod usage time rather than fixed rates, so users pay only for the resources they consume—making it a highly cost-effective option for scalable AI deployments.