The Gemini Enterprise Agent Platform Model Garden SDK helps developers use Model Garden open models to build AI-powered features and applications. The SDKs support use cases like the following:
- Deploy an open model
- Export open model weights
To install the google-cloud-aiplatform Python package, run the following command:
pip3 install --upgrade --user "google-cloud-aiplatform>=1.84"For detailed instructions, see deploy an open model and deploy notebook tutorial.
This is the simplest way to deploy a model. If you provide just a model name, the SDK will use the default deployment configuration.
from agentplatform import model_garden
model = model_garden.OpenModel("google/paligemma@paligemma-224-float32")
endpoint = model.deploy()Use case: Fast prototyping, first-time users evaluating model outputs.
You can list all models that are currently deployable via Model Garden:
from agentplatform import model_garden
models = model_garden.list_deployable_models()To filter only Hugging Face models or by keyword:
models = model_garden.list_deployable_models(list_hf_models=True, model_filter="stable-diffusion")Use case: Discover available models before deciding which one to deploy.
Deploy a model directly from Hugging Face using the model ID.
model = model_garden.OpenModel("Qwen/Qwen2-1.5B-Instruct")
endpoint = model.deploy()Use case: Leverage community or third-party models without custom container setup. If the model is gated, you may need to provide a Hugging Face access token:
endpoint = model.deploy(hugging_face_access_token="your_hf_token")Use case: Deploy gated Hugging Face models requiring authentication.
You can inspect available deployment configurations for a model:
model = model_garden.OpenModel("google/paligemma@paligemma-224-float32")
deploy_options = model.list_deploy_options()Use case: Evaluate compatible machine specs and containers before deployment.
Specify a container image from the list of verified deployment configurations.
endpoint = model.deploy(
serving_container_image_uri="us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250430_0916_RC00_maas",
)Specify a hardware configuration from the list of verified deployment configurations.
endpoints = model.deploy(
machine_type="a3-highgpu-1g",
accelerator_type="NVIDIA_H100_80GB",
accelerator_count=1,
)Specify both a container image and a hardware configuration from the list of verified deployment configurations.
endpoint = model.deploy(
serving_container_image_uri="us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250430_0916_RC00_maas",
machine_type="a3-highgpu-1g",
accelerator_type="NVIDIA_H100_80GB",
accelerator_count=1,
)Use case: Production configuration, performance tuning, scaling.
Some models require acceptance of a license agreement. Pass eula=True if prompted.
model = model_garden.OpenModel("google/gemma2@gemma-2-27b-it")
endpoint = model.deploy(eula=True)Use case: First-time deployment of EULA-protected models.
Schedule workloads on Spot VMs for lower cost.
endpoint = model.deploy(spot=True)Use case: Cost-sensitive development and batch workloads.
Enable experimental fast-deploy path for popular models.
endpoint = model.deploy(fast_tryout_enabled=True)Use case: Interactive experimentation without full production setup.
Create a dedicated DNS-isolated endpoint.
endpoint = model.deploy(use_dedicated_endpoint=True)Use case: Traffic isolation for enterprise or regulated workloads.
Use shared or specific Compute Engine reservations.
endpoint = model.deploy(
reservation_affinity_type="SPECIFIC_RESERVATION",
reservation_affinity_key="compute.googleapis.com/reservation-name",
reservation_affinity_values="projects/YOUR_PROJECT/zones/YOUR_ZONE/reservations/YOUR_RESERVATION"
)Use case: Optimized resource usage with pre-reserved capacity.
Override the default container with a custom image.
endpoint = model.deploy(
serving_container_image_uri="us-docker.pkg.dev/vertex-ai/custom-container:latest"
)Use case: Use of custom inference servers or fine-tuned environments.
Further customize startup probes, health checks, shared memory, and gRPC ports.
endpoint = model.deploy(
serving_container_image_uri="us-docker.pkg.dev/vertex-ai/custom-container:latest",
container_command=["python3"],
container_args=["serve.py"],
container_ports=[8888],
container_env_vars={"ENV": "prod"},
container_predict_route="/predict",
container_health_route="/health",
serving_container_shared_memory_size_mb=512,
serving_container_grpc_ports=[9000],
serving_container_startup_probe_exec=["/bin/check-start.sh"],
serving_container_health_probe_exec=["/bin/health-check.sh"]
)Use case: Production-grade deployments requiring deep customization of runtime behavior and monitoring.
See Contributing for more information on contributing to the Gemini Enterprise Agent Platform Python SDK.
The contents of this repository are licensed under the Apache License, version 2.0.