Skip to main content

AI & assistant-friendly summary

This section provides structured content for AI assistants and search engines. You can cite or summarize it when referencing this page.

Summary

Six reads before you delete a crash-looping pod: context, namespace, get, describe, logs, and events. kubectl client v1.36.1 on 11 October 2026.

Key Facts

  • •kubectl client v1.36.1 on 11 October 2026
  • •On 11 October 2026 printed and
  • •No cluster API calls were made for this article
  • •It also removes the container you have not read yet, unless a controller recreates it and you still have the previous log
  • •The Kubernetes API concept for dry-run authorization is documented at dry-run authorization

Entity Definitions

Bedrock
Bedrock is an AWS service discussed in this article.
IAM
IAM is an AWS service discussed in this article.
EKS
EKS is an AWS service discussed in this article.
DevOps
DevOps is a cloud computing concept discussed in this article.
Kubernetes
Kubernetes is a development tool discussed in this article.
Docker
Docker is a development tool discussed in this article.

Mastering Kubernetes Commands: kubectl, Troubleshooting, and Production Operations

DevOps & CI/CDPalaniappan P10 min read

Quick summary: Six reads before you delete a crash-looping pod: context, namespace, get, describe, logs, and events. kubectl client v1.36.1 on 11 October 2026.

Key Takeaways

  • kubectl client v1.36.1 on 11 October 2026
  • On 11 October 2026 printed and
  • No cluster API calls were made for this article
  • It also removes the container you have not read yet, unless a controller recreates it and you still have the previous log
  • The Kubernetes API concept for dry-run authorization is documented at dry-run authorization
Operations desk with a laptop showing an abstract cluster of connected nodes under an amber lamp.
Table of Contents

On 11 October 2026 kubectl version --client printed Client Version: v1.36.1 and Kustomize Version: v5.8.1. The lab script reported context=set and did not print the context name. No cluster API calls were made for this article. Mutating examples are from the kubectl reference and the debug tasks, and they are marked as not run here. helm was not on PATH.

Six reads before you delete a crash-looping pod: current context, namespace, kubectl get, kubectl describe, kubectl logs, and Events. Deleting the pod restarts it. It also removes the container you have not read yet, unless a controller recreates it and you still have the previous log.

What broke — A viewer who runs kubectl diff can get a forbidden error even though diff persists nothing. Server-side dry run is authorized as a write. The Kubernetes API concept for dry-run authorization is documented at dry-run authorization. The recovery is kubectl auth can-i, not a cluster-admin binding.

Reproduce this — Run bash examples/architecture-blog-2026/mastering-developer-tools/mastering-kubernetes-commands/check-client.sh. Expected: a client version and context=set or context=unset, then lab=ok. It does not call the API server. Published copy: /examples/architecture-blog-2026/mastering-developer-tools/mastering-kubernetes-commands/check-client.sh.

We recommend kubectl apply --dry-run=server on a non-production cluster when your role is allowed to patch, and kubectl apply --dry-run=client when you only need to see the object the client would send. The trade-off: server dry-run catches admission and defaulting, and it requires write RBAC. Client dry-run does not. Neither one is the live apply.

Why kubectl still matters

An agent can emit a manifest that looks valid and schedules nowhere, or it can kubectl delete a Deployment because a pod was unhealthy. Context and namespace decide which cluster that delete hits. A kubeconfig file often contains production and staging in the same file.

Install and identity of the client

Install kubectl from the Kubernetes release that matches your cluster’s minor version as closely as you can. The kubectl install docs are per OS. This page’s client is v1.36.1. A client that is many minors ahead of the API server will warn. Read the warning.

Context: client only. These two were run. Later cluster commands were not.

kubectl version --client
kubectl config get-contexts
kubectl config current-context
kubectl cluster-info

current-context prints a name. The lab hides it on purpose. You should print it before you change anything, and you should recognize production when you see it.

kubectl config set-context --current --namespace=NAMESPACE changes the default namespace for this context. Local change to the kubeconfig. It does not change the cluster. Forgetting you set it is how a delete lands in the wrong namespace. Pass -n NAMESPACE on the command when you are tired.

kubectl config view --minify shows the current context’s cluster and user. Redact certificate data and tokens before you share it. A kubeconfig is a credential.

Help:

kubectl --help
kubectl explain pod.spec.containers
kubectl explain deploy.spec.strategy

explain reads OpenAPI from the cluster when you are connected, and it tells you field types. It is the command to use when an agent invents a field. If explain says the field is absent, do not apply the manifest.

kubectl api-resources lists what this cluster has installed, including CRDs. Gateway API resources appear there only when the CRDs are installed. Do not assume gateway exists on every cluster.

Reads, selectors, and output

TaskCommandRisk
List podskubectl get pods -n NAMESPACERead-only
Wide listkubectl get pods -o wideRead-only
Labelskubectl get pods -l app=catalogRead-only
Sortkubectl get pods --sort-by=.status.startTimeRead-only
YAMLkubectl get pod POD -o yamlRead-only. Can include env from the spec.
JSONPathkubectl get pod POD -o jsonpath='{.status.phase}'Read-only
Describekubectl describe pod PODRead-only. Events at the bottom.
Eventskubectl get events -n NAMESPACE --sort-by=.lastTimestampRead-only

describe is where Pending, image pull errors, and probe failures show up as Events. get alone hides them.

Secrets printed as YAML are secret disclosure. kubectl get secret NAME -o yaml shows data in base64, which is not encryption. Prefer kubectl get secret NAME -o jsonpath='{.metadata.name}' when you only need to know it exists.

Workloads

KindReadChangeRisk of the change
Podget, describeusually created by a controllerDeleting a bare pod removes it.
Deploymentget deploy, rollout statusscale, rollout restart, set imageRemote mutation
ReplicaSetget rsowned by the DeploymentDo not edit old ReplicaSets by hand.
StatefulSetget stsscale is orderedRemote mutation. Identity and storage stick.
DaemonSetget dsone pod per selected nodeRemote mutation
Jobget jobcreate job runs workPotential cost on the cluster
CronJobget cronjobcreate job --from=cronjob/NAMERuns a job now.

kubectl rollout history deployment/NAME and kubectl rollout undo deployment/NAME are the pair for a bad catalog deploy. Undo is remote mutation. Read history first. rollout restart recycles pods and does not change the image. Use it when you need a fresh process, not when the image is wrong.

kubectl scale deployment/NAME --replicas=3 changes desired count. Remote mutation and potential cost on a metered cluster. kubectl autoscale creates an HPA. Read kubectl get hpa before you create a second one.

Network and storage

kubectl get svc,endpoints,ingress -n NAMESPACE shows whether a Service has backends. An endpoints or endpointslice object with no addresses means the selector matches nothing. Pods can be Running and still receive no traffic.

kubectl port-forward pod/POD 8080:8080 opens a local tunnel. It is a local change to your machine’s network and a read or write against the pod port, depending on what you send. Close it when you are done. Do not port-forward a production database to a laptop and leave it.

kubectl cp copies files in or out. Copying a secret file out is disclosure. Copying a binary in is a mutation of the container filesystem that disappears when the pod is replaced. Prefer a new image.

kubectl get pv,pvc shows binds. A Pending pod that says unbound PVC is a storage problem. describe pvc names the StorageClass and the event. Deleting a PVC can delete the volume, depending on the reclaim policy. Potentially destructive. Read describe pv for persistentVolumeReclaimPolicy before you delete.

Apply, diff, and dry run

kubectl diff -f manifest.yaml shows a diff and uses server-side dry run. It needs permission to patch or create. It should not persist the object. “Should not” is the API’s dry-run contract. A broken webhook can still make diff fail. Read the error.

kubectl apply --dry-run=client -f manifest.yaml renders the object locally and does not contact the mutating path the same way. kubectl apply --dry-run=server -f manifest.yaml asks the API server to admit the object and return it without saving. kubectl apply -f manifest.yaml without dry-run writes. Remote mutation.

kubectl apply --server-side -f manifest.yaml is server-side apply, which is a real write with field management. It is not the same flag as --dry-run=server. Read both words before you approve an agent line.

kubectl patch and kubectl set edit live objects. Record kubectl get -o yaml first if you do not have GitOps. GitOps for EKS is covered in EKS with Argo CD and Flux.

kubectl rollout status deployment/NAME waits until the deployment reports success or you interrupt it. A command that returns 0 means the rollout controller is satisfied. Check the app with a request or a probe, not only with that exit code.

Logs, exec, and previous containers

kubectl logs POD -n NAMESPACE --tail=100
kubectl logs POD -c CONTAINER --previous
kubectl logs deployment/NAME --tail=50

--previous is the log of the last dead container. That is the crash-loop evidence. Without it you see the new container, which may not have failed yet. These were not run here. Confirm the flag with kubectl logs --help on v1.36.

kubectl exec POD -- COMMAND runs a command in the container. It is not read-only. A shell as root inside the container is a production change if the process can write.

Debug pods and ephemeral containers exist in current Kubernetes. kubectl debug is the documented entry. Read kubectl debug --help before you attach a privileged debug container to a production node. A privileged debug container is potentially destructive and a security event.

RBAC

kubectl auth can-i get pods -n NAMESPACE
kubectl auth can-i create deployments.apps -n NAMESPACE
kubectl auth can-i --list -n NAMESPACE

can-i is read-only against the authorization API. --list can be long. It tells you what this kubeconfig can do. If the answer is no, do not bind cluster-admin to finish a ticket.

kubectl get role,rolebinding,clusterrolebinding -n NAMESPACE shows what was granted. ClusterRoleBindings are cluster-scoped. A binding you did not mean to create is potentially destructive because of the access it grants. Delete it only after you know which workload uses it.

Service account tokens in Secrets or projected volumes are credentials. Do not kubectl get secret -o yaml them into a chat with an agent.

Scenario: an order worker in CrashLoopBackOff

On a non-production cluster first if you have one. If you must look at production, stay on the six reads.

  1. kubectl config current-context and kubectl config view --minify until you can say the cluster name out loud.
  2. kubectl get pods -n NAMESPACE -l app=order-worker. Note Ready, Status, and Restarts.
  3. kubectl describe pod POD. Read the Events. ImagePullBackOff, insufficient cpu, and failed probes are different faults.
  4. kubectl logs POD --previous --tail=200. OOMKilled shows up here and in describe as a last state reason.
  5. kubectl top pod POD only if metrics-server is installed. If top errors, say metrics are missing. Do not install cluster extras in the middle of an incident unless that is the change you planned.
  6. kubectl rollout history deployment/NAME if a deploy just went out. Compare the image with the previous revision before rollout undo.

Pending is not CrashLoopBackOff. Pending means the pod has not started: scheduler, PVC, or taint. Describe tells you which. CrashLoopBackOff means the container process exits and Kubernetes starts it again. Deleting the pod repeats the cycle if the Deployment is unchanged.

Image pull failures: the event names the image and the registry error. Fix the pull secret or the tag. Do not delete the node.

Probe failures: the pod may be Running and not Ready. Traffic should not arrive until Ready. Editing the probe to always succeed hides a broken process.

Nodes and maintenance

kubectl get nodes and kubectl describe node NODE show pressure and taints. Read-only.

kubectl cordon NODE marks the node unschedulable. Remote mutation. Existing pods stay. kubectl drain NODE evicts pods that can be evicted. Remote mutation that causes downtime if the workload has no spare replicas or if a PodDisruptionBudget blocks the eviction and you override it. Drain production nodes only with a reason, a PDB, and a way back (kubectl uncordon).

kubectl delete pod on a Deployment recreates the pod. kubectl delete deployment removes the workload. Potentially destructive. kubectl delete -f manifest.yaml deletes what the file selects. Read the file.

Force deletion (--force --grace-period=0) is for API objects stuck on a dead node. It is not the first response to CrashLoopBackOff.

Helm, optional and not run here

Helm was not installed on the authoring machine on 11 October 2026. Do not treat this list as a tested Helm session. The usual reads, from Helm’s own docs, are helm list -n NAMESPACE, helm status RELEASE -n NAMESPACE, and helm history RELEASE -n NAMESPACE. helm upgrade --dry-run renders and, depending on flags and version, may still talk to the cluster. Run helm upgrade --help on the binary you have before a production upgrade. helm uninstall deletes the release. Potentially destructive.

kubectl remains the tool that shows the pods Helm created. If Helm says deployed and the pods crash, you are back in the six reads.

Working with a coding agent

Ask the agent to print current-context and the namespace before any write. Ask for diff or --dry-run=server and read the error if RBAC blocks it. Refuse delete, drain, and cluster-admin bindings until the six reads are in the ticket. After a rollout, kubectl rollout status and a log tail are the check. The agent saying the manifest applied is the start of that check.

Container filesystem questions that are really image questions belong in Docker. IAM for the node role belongs in AWS CLI.

Five labs

Use a cluster you can break: kind, k3d, or a sandbox namespace. Do not use production.

  1. Run check-client.sh. If context=set, run kubectl config current-context yourself and decide whether that cluster is fair game.
  2. kubectl get ns and kubectl auth can-i get pods -n default. Expected: yes or no. Either answer is the lesson.
  3. Apply a Deployment of a public image you trust with --dry-run=server if can-i says you may create deployments. Read the output. Then apply for real only in the sandbox.
  4. Point the container command at exit 1. Watch CrashLoopBackOff with describe and logs --previous. Fix the command. Do not delete the Deployment as the first move.
  5. kubectl rollout history and kubectl rollout undo on that sandbox Deployment. Expected: the previous pod template returns. Confirm with kubectl get pods.

Progression: context and can-i, then get and describe, then diff and server dry-run, then rollout undo, and only then drain or delete.

What this post does not cover

Cluster build (EKS version upgrades, Karpenter), service mesh, and policy engines. Runtime seccomp and AppArmor are in the container runtime security post. This page will go stale on flag details. kubectl COMMAND --help wins.

What to do this week

  1. Run the client check and write down the context you use for staging.
  2. Practice logs --previous on a sandbox crash loop once, before a real order worker does it.
  3. Remove any personal kubeconfig entry you no longer need. The file is a credential.
  4. If deploys go through Git, read the GitOps note and keep kubectl for diagnosis.

Quick reference

I need toCommandRisk
Know the clusterkubectl config current-contextRead-only
See why a pod waitskubectl describe podRead-only
See the crashed processkubectl logs --previousRead-only
Preview a writekubectl apply --dry-run=serverNeeds write RBAC. Should not persist.
Undo a Deploymentkubectl rollout undoRemote mutation
Evict a nodekubectl drainRemote mutation

You should be able to say which cluster you are on, separate Pending from CrashLoopBackOff, and refuse cluster-admin as a debugging step.

Further reading

Contact us or see DevOps pipeline setup if production kubectl access is shared from one kubeconfig and nobody owns the RBAC review.

Frequently asked questions

When should you not grant cluster-admin to debug a pod?
As a routine fix. cluster-admin can read every secret and delete every namespace. Use kubectl auth can-i to see the missing verb, then ask for that verb on that resource in that namespace.
Does kubectl diff only read the cluster?
kubectl diff submits a server-side dry run. The API server authorizes it like a write: patch on an existing object, create on a new one. A viewer role can fail diff even though nothing is persisted. Client-side dry-run (kubectl apply --dry-run=client) does not send that write.
What is the difference between delete and force delete?
A normal delete asks the pod to shut down and respects the grace period. Force delete with a zero grace period removes the API object even if the kubelet has not confirmed the process is gone. Use it only when the node is dead and you accept that a process might still be running there.
Helm and kubectl are the same tool?
No. kubectl talks to the Kubernetes API. Helm renders charts and stores release state. This page's authoring machine did not have helm on PATH. Treat Helm commands as a separate optional section and run helm --help before you rely on a flag.
Does this page list every kubectl command?
No. The full tree is kubectl --help and the Kubernetes kubectl reference. This page is the operations set for workloads, network, auth, and recovery.
Palaniappan P
Palaniappan P

AWS Cloud Architect & AI Expert

AWS-certified cloud architect and AI expert with deep expertise in cloud migrations, cost optimization, and generative AI on AWS.

AWS ArchitectureCloud MigrationGenAI on AWSCost OptimizationDevOps

Recommended Reading

Explore All Articles »
5 min

GitOps on Amazon EKS (2026): Argo CD vs Flux, App-of-Apps, and the Decisions That Actually Bite

AWS Prescriptive Guidance says Argo CD and Flux both handle most GitOps scenarios capably — so picking one is a fit decision, not a winner. The decisions that actually cause incidents are the ones underneath: plaintext secrets in the GitOps repo, CI running kubectl apply and reintroducing drift, no App-of-Apps so onboarding is click-ops, and repo topology you can't change later. Here is the Argo CD vs Flux matrix, an App-of-Apps example, and the five traps independent of tool.