Catalog
google/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring

google

gke-ai-troubleshooting-tpu-dynamic-slices-monitoring

Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Use when checking TPU slice lifecycle states, troubleshooting slice provisioning failures, validating single-slice or multi-slice (JobSet) workload manifests, or safely patching stuck finalizers and disabling the slice controller. Don't use for generic GKE cluster node pool creation or standard non-TPU workload management (use gke-basics or gke-cluster-creation instead).

v1.0LATEST
New~1.9kUpdated Aug 31, 2026

GKE TPU Dynamic Slices Monitoring & Management

Monitors the status of TPU Slice custom resources, troubleshoots provisioning failures, validates workload manifests on dynamic slices, and performs cleanups.

Prerequisites

  • Cloud Logging enabled for the project.
  • kubectl and gcloud CLIs configured to access the GKE cluster.

Diagnostic Workflow

Step 0: Context Acquisition & Time Window Definition

Gather project, cluster, and slice context using cluster tools or the following parameters:

  • Project ID: {project_id} (e.g., my-gcp-project)
  • Cluster Name: {cluster_name} (e.g., tpu-cluster)
  • Region/Zone: {location} (e.g., us-central1-a)
  • Slice Name: {slice_name} (e.g., test-slice)
  • Issue Time: {timestamp} (Optional; default to the last 30 minutes window [T - 30m] to [T + 30m])

Step 1: Describe the Slice Custom Resource [Low Risk]

When asked to inspect, troubleshoot, or check a slice status, immediately execute kubectl describe slice {slice_name} using available cluster tools to perform the inspection. Parse the resulting Status.Conditions output against the condition table below to diagnose the exact state and provide concrete recommendations.

  • Command:

    kubectl describe slice {slice_name}
    

State & Reason Analysis

Analyze the Status.Conditions (especially Type: Ready and its Reason and Status):

Lifecycle State / Reason Meaning Recommended Action
SliceNotCreated GKE Slice Controller is initializing the slice and performing resource checks. Wait a few minutes and re-check slice status.
SliceCreationFailed Prerequisites validation failed (e.g., selected nodes don't exist, nodes are already used by another slice, or the topology doesn't match the number of partitions). Verify selected nodes exist, are unallocated, and topology matches partition count.
ACTIVATING GKE is actively forming and provisioning the TPU slice. Monitor node provisioning.
ACTIVE The TPU slice is successfully formed and ready to host workloads. Proceed to deploy or check workloads.
ACTIVE_DEGRADED The slice is usable, but one or more sub-blocks are degraded. Monitor workload logs for interconnect or device errors. Check faulty node VMs.
FAILED GKE failed to form the TPU slice (e.g., selected nodes are not part of the same reservation block). Ensure all selected nodes belong to the same reservation block.
DEACTIVATING The slice is dismantling (triggered by user deletion or a critical systemic failure). Wait for dismantling to finish, or patch finalizers if stuck.
INCOMPLETE The terminal phase before the Slice CR is deleted from the cluster. No action required; the resource will be removed shortly.

Provisioning Failure Troubleshooting Checklist

When investigating slice creation or provisioning failures (SliceCreationFailed or FAILED), perform the following verification steps:

  1. Node Existence & Allocation Check: Verify that the selected TPU nodes exist in the cluster and are not already allocated to another slice (kubectl get nodes -l cloud.google.com/gke-tpu-slice, kubectl get slice -A).
  2. Topology Alignment: Confirm that the partition count matches the requested topology dimensions (e.g. topology 2x2 requires 4 nodes).
  3. Reservation Block Alignment Check: Confirm that all selected TPU nodes belong to the same reservation and reservation block.

Step 2: Verify Workload Specification [Low Risk]

Ensure workload manifests are configured correctly to target the dynamic slice.

1. Single-Slice Workload Requirements

Check that the Pod template contains the following annotations and selectors:

  • Annotations:
    • cloud.google.com/gke-tpu-slice-topology: "{topology}" (e.g., "4x4x4")
  • NodeSelector:
    • cloud.google.com/gke-tpu-topology: "{topology}" (e.g., "4x4x4")
    • cloud.google.com/gke-tpu-accelerator: "{accelerator_type}" (e.g., "tpu7x")
    • cloud.google.com/gke-tpu-slice: "{slice_name}" (e.g., "test-slice")

2. Multi-Slice (JobSet) Workload Requirements

If deploying a multi-slice JobSet, verify:

  • JobSet Annotation:
    • alpha.jobset.sigs.k8s.io/exclusive-topology: cloud.google.com/gke-tpu-slice
  • Pod Template Annotations:
    • cloud.google.com/gke-tpu-slice-topology: "{topology}"
  • Pod Template NodeSelector:
    • cloud.google.com/gke-tpu-topology: "{topology}"
    • cloud.google.com/gke-tpu-accelerator: "{accelerator_type}"
    • Note: Do NOT manually specify cloud.google.com/gke-tpu-slice in the nodeSelector; JobSet handles slice assignment automatically.

Resolution & Management Workflow

Resolution 1: Force Delete a Stuck Slice [High Risk]

If a slice is stuck in DEACTIVATING or deletion hangs indefinitely due to stuck finalizers:

  1. Identify Cause: Explain that finalizers on the slice resource (metadata.finalizers) are preventing Kubernetes from completing resource deletion.

  2. Propose Resolution: Propose removing finalizers from the metadata path (/metadata/finalizers) using a JSON patch operation:

    kubectl patch slice {slice_name} --type json -p='[{"op": "remove", "path": "/metadata/finalizers"}]'
    
  3. Provide Warning: Explicitly warn the user that removing finalizers bypasses standard controller dismantling and may leave underlying VM, network, or accelerator resources uncleaned or orphaned.

  4. CRITICAL SAFETY MANDATE: The response MUST explicitly ask the user for confirmation (e.g. "Removing finalizers on /metadata/finalizers via JSON patch is a high-risk operation that may leave orphaned resources. Do you confirm you want to apply this patch to slice {slice_name}?") and pause for user confirmation before applying or executing the patch.


Resolution 2: Disable and Clean Up Slice Controller [High Risk]

If dynamic slicing needs to be disabled:

  1. Check for existing Slices:

    kubectl get slice -A
    

    Ensure all slices are deleted before disabling the controller.

  2. Disable Slice Controller via gcloud:

    gcloud container clusters update {cluster_name} \
        --location={location} \
        --no-enable-slice-controller
    
  3. Delete the Slice CRD:

    kubectl delete crd slices.accelerator.gke.io
    
  4. Clean up Node Labels: Remove GKE TPU Slice labels from all nodes in the cluster:

    kubectl label nodes --all cloud.google.com/gke-tpu-slice- cloud.google.com/gke-tpu-slice-topology-
    
  • Safety Rule: Propose the exact commands and confirm before executing disabling or destructive cleanup steps.
Files3
3 files · 13.2 KB

Select a file to preview

Overall Score

82/100

Grade

B

Good

Safety

80

Quality

85

Clarity

82

Completeness

78

Summary

This skill provides structured guidance for monitoring, troubleshooting, and managing GKE TPU Dynamic Slices custom resources. It covers slice lifecycle diagnostics, workload manifest validation for single and multi-slice JobSet deployments, and two high-risk administrative operations: force-deleting stuck slices via finalizer removal and disabling the slice controller. The skill is narrowly scoped to TPU dynamic slice operations and explicitly excludes generic GKE node pool or non-TPU workload management.

Detected Capabilities

kubectl executionkubectl describe (read-only inspection)kubectl patch (metadata modification)kubectl delete (resource deletion)gcloud CLI executionnode label removalCRD deletion

Trigger Keywords

Phrases that MCP clients use to match this skill to user intent.

tpu slice troubleshootinggke tpu dynamic slicestpu slice provisioningtpu slice stuck finalizertpu slice statusjobset tpu validation

Risk Signals

WARNING

kubectl patch with JSON finalizer removal operation

Resolution 1: Force Delete a Stuck Slice section
WARNING

Disabling slice controller via gcloud and deleting CRD (slices.accelerator.gke.io)

Resolution 2: Disable and Clean Up Slice Controller section
WARNING

Cluster-wide node label removal using 'kubectl label nodes --all'

Resolution 2: Disable and Clean Up Slice Controller section

Referenced Domains

External domains referenced in skill content, detected by static analysis.

www.apache.org

Use Cases

  • Monitor TPU slice provisioning status and lifecycle transitions
  • Troubleshoot TPU slice creation failures and validation errors
  • Validate single-slice Pod manifests targeting dynamic slices
  • Validate multi-slice JobSet workload configurations
  • Force-delete TPU slices stuck in DEACTIVATING state due to hung finalizers
  • Disable GKE TPU slice controller and clean up related resources

Quality Notes

  • Clear boundary definition in description: scope is limited to TPU dynamic slices, explicitly excludes generic GKE operations
  • Comprehensive state/reason analysis table maps all slice lifecycle conditions to diagnostic meanings and recommended actions
  • Well-structured diagnostic workflow with numbered steps and clear context acquisition parameters
  • Detailed checklists for troubleshooting provisioning failures (node existence, topology alignment, reservation block checks)
  • Both single-slice and multi-slice (JobSet) workload validation requirements are documented with precise annotation and nodeSelector specifications
  • High-risk operations are properly marked and include explicit user confirmation requirements
  • Supporting reference file (failure_signatures.md) provides concrete examples of YAML conditions for realistic diagnostics
  • All destructive/high-risk operations include safety warnings and propose commands before execution
  • Validation script validates Cloud Logging queries (though none are defined in this skill)
  • Limitations are clear: do not use for generic GKE node pool creation or standard non-TPU workload management
Model: claude-haiku-4-5-20251001Analyzed: Aug 31, 2026

Reviews

Add this skill to your library to leave a review.

No reviews yet

Be the first to share your experience.

Use google/gke-ai-troubleshooting-tpu-dynamic-slices-monitoring in your dev environment

Command Palette

Search for a command to run...