Skip to content

Latest commit

 

History

History
 
 

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 

README.md

Gemini Enterprise Agent Platform Model Garden SDK for Python

The Gemini Enterprise Agent Platform Model Garden SDK helps developers use Model Garden open models to build AI-powered features and applications. The SDKs support use cases like the following:

  • Deploy an open model
  • Export open model weights

Installation

To install the google-cloud-aiplatform Python package, run the following command:

pip3 install --upgrade --user "google-cloud-aiplatform>=1.84"

Usage

For detailed instructions, see deploy an open model and deploy notebook tutorial.

Quick Start: Default Deployment

This is the simplest way to deploy a model. If you provide just a model name, the SDK will use the default deployment configuration.

from agentplatform import model_garden

model = model_garden.OpenModel("google/paligemma@paligemma-224-float32")
endpoint = model.deploy()

Use case: Fast prototyping, first-time users evaluating model outputs.

List Deployable Models

You can list all models that are currently deployable via Model Garden:

from agentplatform import model_garden

models = model_garden.list_deployable_models()

To filter only Hugging Face models or by keyword:

models = model_garden.list_deployable_models(list_hf_models=True, model_filter="stable-diffusion")

Use case: Discover available models before deciding which one to deploy.

Hugging Face Model Deployment

Deploy a model directly from Hugging Face using the model ID.

model = model_garden.OpenModel("Qwen/Qwen2-1.5B-Instruct")
endpoint = model.deploy()

Use case: Leverage community or third-party models without custom container setup. If the model is gated, you may need to provide a Hugging Face access token:

endpoint = model.deploy(hugging_face_access_token="your_hf_token")

Use case: Deploy gated Hugging Face models requiring authentication.

List Deployment Configurations

You can inspect available deployment configurations for a model:

model = model_garden.OpenModel("google/paligemma@paligemma-224-float32")
deploy_options = model.list_deploy_options()

Use case: Evaluate compatible machine specs and containers before deployment.

Select a Verified Deployment: By Container Image

Specify a container image from the list of verified deployment configurations.

endpoint = model.deploy(
    serving_container_image_uri="us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250430_0916_RC00_maas",
)

Select a Verified Deployment: By Hardware

Specify a hardware configuration from the list of verified deployment configurations.

endpoints = model.deploy(
    machine_type="a3-highgpu-1g",
    accelerator_type="NVIDIA_H100_80GB",
    accelerator_count=1,
)

Select a Verified Deployment: By Container and Hardware

Specify both a container image and a hardware configuration from the list of verified deployment configurations.

endpoint = model.deploy(
    serving_container_image_uri="us-docker.pkg.dev/vertex-ai/vertex-vision-model-garden-dockers/pytorch-vllm-serve:20250430_0916_RC00_maas",
    machine_type="a3-highgpu-1g",
    accelerator_type="NVIDIA_H100_80GB",
    accelerator_count=1,
)

Use case: Production configuration, performance tuning, scaling.

EULA Acceptance

Some models require acceptance of a license agreement. Pass eula=True if prompted.

model = model_garden.OpenModel("google/gemma2@gemma-2-27b-it")
endpoint = model.deploy(eula=True)

Use case: First-time deployment of EULA-protected models.

Spot VM Deployment

Schedule workloads on Spot VMs for lower cost.

endpoint = model.deploy(spot=True)

Use case: Cost-sensitive development and batch workloads.

Fast Tryout Deployment

Enable experimental fast-deploy path for popular models.

endpoint = model.deploy(fast_tryout_enabled=True)

Use case: Interactive experimentation without full production setup.

Dedicated Endpoints

Create a dedicated DNS-isolated endpoint.

endpoint = model.deploy(use_dedicated_endpoint=True)

Use case: Traffic isolation for enterprise or regulated workloads.

Reservation Affinity

Use shared or specific Compute Engine reservations.

endpoint = model.deploy(
    reservation_affinity_type="SPECIFIC_RESERVATION",
    reservation_affinity_key="compute.googleapis.com/reservation-name",
    reservation_affinity_values="projects/YOUR_PROJECT/zones/YOUR_ZONE/reservations/YOUR_RESERVATION"
)

Use case: Optimized resource usage with pre-reserved capacity.

Custom Container Image

Override the default container with a custom image.

endpoint = model.deploy(
    serving_container_image_uri="us-docker.pkg.dev/vertex-ai/custom-container:latest"
)

Use case: Use of custom inference servers or fine-tuned environments.

Advanced Full Container Configuration

Further customize startup probes, health checks, shared memory, and gRPC ports.

endpoint = model.deploy(
    serving_container_image_uri="us-docker.pkg.dev/vertex-ai/custom-container:latest",
    container_command=["python3"],
    container_args=["serve.py"],
    container_ports=[8888],
    container_env_vars={"ENV": "prod"},
    container_predict_route="/predict",
    container_health_route="/health",
    serving_container_shared_memory_size_mb=512,
    serving_container_grpc_ports=[9000],
    serving_container_startup_probe_exec=["/bin/check-start.sh"],
    serving_container_health_probe_exec=["/bin/health-check.sh"]
)

Use case: Production-grade deployments requiring deep customization of runtime behavior and monitoring.

Contributing

See Contributing for more information on contributing to the Gemini Enterprise Agent Platform Python SDK.

License

The contents of this repository are licensed under the Apache License, version 2.0.