---
title: Mastering Kubernetes Commands: kubectl, Troubleshooting, and Production Operations
description: Six reads before you delete a crash-looping pod: context, namespace, get, describe, logs, and events. kubectl client v1.36.1 on 11 October 2026.
url: https://www.factualminds.com/blog/mastering-kubernetes-commands/
datePublished: 2026-10-11T00:00:00.000Z
dateModified: 2026-10-11T00:00:00.000Z
author: palaniappan-p
category: DevOps & CI/CD
tags: kubernetes, kubectl, eks, devops
---

# Mastering Kubernetes Commands: kubectl, Troubleshooting, and Production Operations

> Six reads before you delete a crash-looping pod: context, namespace, get, describe, logs, and events. kubectl client v1.36.1 on 11 October 2026.

On 11 October 2026 `kubectl version --client` printed `Client Version: v1.36.1` and `Kustomize Version: v5.8.1`. The lab script reported `context=set` and did not print the context name. No cluster API calls were made for this article. Mutating examples are from the [kubectl reference](https://kubernetes.io/docs/reference/kubectl/) and the [debug tasks](https://kubernetes.io/docs/tasks/debug/), and they are marked as not run here. `helm` was not on `PATH`.

Six reads before you delete a crash-looping pod: current context, namespace, `kubectl get`, `kubectl describe`, `kubectl logs`, and Events. Deleting the pod restarts it. It also removes the container you have not read yet, unless a controller recreates it and you still have the previous log.

> **What broke** — A viewer who runs `kubectl diff` can get a forbidden error even though diff persists nothing. Server-side dry run is authorized as a write. The Kubernetes API concept for dry-run authorization is documented at [dry-run authorization](https://kubernetes.io/docs/reference/using-api/api-concepts/#dry-run-authorization). The recovery is `kubectl auth can-i`, not a cluster-admin binding.

> **Reproduce this** — Run `bash examples/architecture-blog-2026/mastering-developer-tools/mastering-kubernetes-commands/check-client.sh`. Expected: a client version and `context=set` or `context=unset`, then `lab=ok`. It does not call the API server. Published copy: [/examples/architecture-blog-2026/mastering-developer-tools/mastering-kubernetes-commands/check-client.sh](/examples/architecture-blog-2026/mastering-developer-tools/mastering-kubernetes-commands/check-client.sh).

We recommend `kubectl apply --dry-run=server` on a non-production cluster when your role is allowed to patch, and `kubectl apply --dry-run=client` when you only need to see the object the client would send. The trade-off: server dry-run catches admission and defaulting, and it requires write RBAC. Client dry-run does not. Neither one is the live apply.

## Why kubectl still matters

An agent can emit a manifest that looks valid and schedules nowhere, or it can `kubectl delete` a Deployment because a pod was unhealthy. Context and namespace decide which cluster that delete hits. A kubeconfig file often contains production and staging in the same file.

## Install and identity of the client

Install kubectl from the Kubernetes release that matches your cluster's minor version as closely as you can. The [kubectl install docs](https://kubernetes.io/docs/tasks/tools/) are per OS. This page's client is v1.36.1. A client that is many minors ahead of the API server will warn. Read the warning.

Context: client only. These two were run. Later cluster commands were not.

```bash
kubectl version --client
kubectl config get-contexts
kubectl config current-context
kubectl cluster-info
```

`current-context` prints a name. The lab hides it on purpose. You should print it before you change anything, and you should recognize production when you see it.

`kubectl config set-context --current --namespace=NAMESPACE` changes the default namespace for this context. **Local change** to the kubeconfig. It does not change the cluster. Forgetting you set it is how a delete lands in the wrong namespace. Pass `-n NAMESPACE` on the command when you are tired.

`kubectl config view --minify` shows the current context's cluster and user. Redact certificate data and tokens before you share it. A kubeconfig is a credential.

Help:

```bash
kubectl --help
kubectl explain pod.spec.containers
kubectl explain deploy.spec.strategy
```

`explain` reads OpenAPI from the cluster when you are connected, and it tells you field types. It is the command to use when an agent invents a field. If explain says the field is absent, do not apply the manifest.

`kubectl api-resources` lists what this cluster has installed, including CRDs. Gateway API resources appear there only when the CRDs are installed. Do not assume `gateway` exists on every cluster.

## Reads, selectors, and output

| Task | Command | Risk |
| --- | --- | --- |
| List pods | `kubectl get pods -n NAMESPACE` | **Read-only** |
| Wide list | `kubectl get pods -o wide` | **Read-only** |
| Labels | `kubectl get pods -l app=catalog` | **Read-only** |
| Sort | `kubectl get pods --sort-by=.status.startTime` | **Read-only** |
| YAML | `kubectl get pod POD -o yaml` | **Read-only**. Can include env from the spec. |
| JSONPath | `kubectl get pod POD -o jsonpath='{.status.phase}'` | **Read-only** |
| Describe | `kubectl describe pod POD` | **Read-only**. Events at the bottom. |
| Events | `kubectl get events -n NAMESPACE --sort-by=.lastTimestamp` | **Read-only** |

`describe` is where Pending, image pull errors, and probe failures show up as Events. `get` alone hides them.

Secrets printed as YAML are **secret disclosure**. `kubectl get secret NAME -o yaml` shows data in base64, which is not encryption. Prefer `kubectl get secret NAME -o jsonpath='{.metadata.name}'` when you only need to know it exists.

## Workloads

| Kind | Read | Change | Risk of the change |
| --- | --- | --- | --- |
| Pod | `get`, `describe` | usually created by a controller | Deleting a bare pod removes it. |
| Deployment | `get deploy`, `rollout status` | `scale`, `rollout restart`, `set image` | **Remote mutation** |
| ReplicaSet | `get rs` | owned by the Deployment | Do not edit old ReplicaSets by hand. |
| StatefulSet | `get sts` | `scale` is ordered | **Remote mutation**. Identity and storage stick. |
| DaemonSet | `get ds` | one pod per selected node | **Remote mutation** |
| Job | `get job` | `create job` runs work | **Potential cost** on the cluster |
| CronJob | `get cronjob` | `create job --from=cronjob/NAME` | Runs a job now. |

`kubectl rollout history deployment/NAME` and `kubectl rollout undo deployment/NAME` are the pair for a bad catalog deploy. Undo is **remote mutation**. Read history first. `rollout restart` recycles pods and does not change the image. Use it when you need a fresh process, not when the image is wrong.

`kubectl scale deployment/NAME --replicas=3` changes desired count. **Remote mutation** and **potential cost** on a metered cluster. `kubectl autoscale` creates an HPA. Read `kubectl get hpa` before you create a second one.

## Network and storage

`kubectl get svc,endpoints,ingress -n NAMESPACE` shows whether a Service has backends. An endpoints or endpointslice object with no addresses means the selector matches nothing. Pods can be Running and still receive no traffic.

`kubectl port-forward pod/POD 8080:8080` opens a local tunnel. It is a **local change** to your machine's network and a read or write against the pod port, depending on what you send. Close it when you are done. Do not port-forward a production database to a laptop and leave it.

`kubectl cp` copies files in or out. Copying a secret file out is disclosure. Copying a binary in is a mutation of the container filesystem that disappears when the pod is replaced. Prefer a new image.

`kubectl get pv,pvc` shows binds. A Pending pod that says unbound PVC is a storage problem. `describe pvc` names the StorageClass and the event. Deleting a PVC can delete the volume, depending on the reclaim policy. **Potentially destructive.** Read `describe pv` for `persistentVolumeReclaimPolicy` before you delete.

## Apply, diff, and dry run

`kubectl diff -f manifest.yaml` shows a diff and uses server-side dry run. It needs permission to patch or create. It should not persist the object. "Should not" is the API's dry-run contract. A broken webhook can still make diff fail. Read the error.

`kubectl apply --dry-run=client -f manifest.yaml` renders the object locally and does not contact the mutating path the same way. `kubectl apply --dry-run=server -f manifest.yaml` asks the API server to admit the object and return it without saving. `kubectl apply -f manifest.yaml` without dry-run writes. **Remote mutation.**

`kubectl apply --server-side -f manifest.yaml` is server-side apply, which is a real write with field management. It is not the same flag as `--dry-run=server`. Read both words before you approve an agent line.

`kubectl patch` and `kubectl set` edit live objects. Record `kubectl get -o yaml` first if you do not have GitOps. GitOps for EKS is covered in [EKS with Argo CD and Flux](/blog/aws-gitops-eks-argocd-flux-2026/).

`kubectl rollout status deployment/NAME` waits until the deployment reports success or you interrupt it. A command that returns 0 means the rollout controller is satisfied. Check the app with a request or a probe, not only with that exit code.

## Logs, exec, and previous containers

```bash
kubectl logs POD -n NAMESPACE --tail=100
kubectl logs POD -c CONTAINER --previous
kubectl logs deployment/NAME --tail=50
```

`--previous` is the log of the last dead container. That is the crash-loop evidence. Without it you see the new container, which may not have failed yet. These were not run here. Confirm the flag with `kubectl logs --help` on v1.36.

`kubectl exec POD -- COMMAND` runs a command in the container. It is not read-only. A shell as root inside the container is a production change if the process can write.

Debug pods and ephemeral containers exist in current Kubernetes. `kubectl debug` is the documented entry. Read `kubectl debug --help` before you attach a privileged debug container to a production node. A privileged debug container is **potentially destructive** and a security event.

## RBAC

```bash
kubectl auth can-i get pods -n NAMESPACE
kubectl auth can-i create deployments.apps -n NAMESPACE
kubectl auth can-i --list -n NAMESPACE
```

`can-i` is **read-only** against the authorization API. `--list` can be long. It tells you what this kubeconfig can do. If the answer is `no`, do not bind `cluster-admin` to finish a ticket.

`kubectl get role,rolebinding,clusterrolebinding -n NAMESPACE` shows what was granted. ClusterRoleBindings are cluster-scoped. A binding you did not mean to create is **potentially destructive** because of the access it grants. Delete it only after you know which workload uses it.

Service account tokens in Secrets or projected volumes are credentials. Do not `kubectl get secret -o yaml` them into a chat with an agent.

## Scenario: an order worker in CrashLoopBackOff

On a non-production cluster first if you have one. If you must look at production, stay on the six reads.

1. `kubectl config current-context` and `kubectl config view --minify` until you can say the cluster name out loud.
2. `kubectl get pods -n NAMESPACE -l app=order-worker`. Note Ready, Status, and Restarts.
3. `kubectl describe pod POD`. Read the Events. ImagePullBackOff, insufficient cpu, and failed probes are different faults.
4. `kubectl logs POD --previous --tail=200`. OOMKilled shows up here and in `describe` as a last state reason.
5. `kubectl top pod POD` only if metrics-server is installed. If `top` errors, say metrics are missing. Do not install cluster extras in the middle of an incident unless that is the change you planned.
6. `kubectl rollout history deployment/NAME` if a deploy just went out. Compare the image with the previous revision before `rollout undo`.

Pending is not CrashLoopBackOff. Pending means the pod has not started: scheduler, PVC, or taint. Describe tells you which. CrashLoopBackOff means the container process exits and Kubernetes starts it again. Deleting the pod repeats the cycle if the Deployment is unchanged.

Image pull failures: the event names the image and the registry error. Fix the pull secret or the tag. Do not delete the node.

Probe failures: the pod may be Running and not Ready. Traffic should not arrive until Ready. Editing the probe to always succeed hides a broken process.

## Nodes and maintenance

`kubectl get nodes` and `kubectl describe node NODE` show pressure and taints. **Read-only.**

`kubectl cordon NODE` marks the node unschedulable. **Remote mutation.** Existing pods stay. `kubectl drain NODE` evicts pods that can be evicted. **Remote mutation** that causes downtime if the workload has no spare replicas or if a PodDisruptionBudget blocks the eviction and you override it. Drain production nodes only with a reason, a PDB, and a way back (`kubectl uncordon`).

`kubectl delete pod` on a Deployment recreates the pod. `kubectl delete deployment` removes the workload. **Potentially destructive.** `kubectl delete -f manifest.yaml` deletes what the file selects. Read the file.

Force deletion (`--force --grace-period=0`) is for API objects stuck on a dead node. It is not the first response to CrashLoopBackOff.

## Helm, optional and not run here

Helm was not installed on the authoring machine on 11 October 2026. Do not treat this list as a tested Helm session. The usual reads, from Helm's own docs, are `helm list -n NAMESPACE`, `helm status RELEASE -n NAMESPACE`, and `helm history RELEASE -n NAMESPACE`. `helm upgrade --dry-run` renders and, depending on flags and version, may still talk to the cluster. Run `helm upgrade --help` on the binary you have before a production upgrade. `helm uninstall` deletes the release. **Potentially destructive.**

kubectl remains the tool that shows the pods Helm created. If Helm says deployed and the pods crash, you are back in the six reads.

## Working with a coding agent

Ask the agent to print `current-context` and the namespace before any write. Ask for `diff` or `--dry-run=server` and read the error if RBAC blocks it. Refuse `delete`, `drain`, and `cluster-admin` bindings until the six reads are in the ticket. After a rollout, `kubectl rollout status` and a log tail are the check. The agent saying the manifest applied is the start of that check.

Container filesystem questions that are really image questions belong in [Docker](/blog/mastering-docker-commands/). IAM for the node role belongs in [AWS CLI](/blog/mastering-aws-cli/).

## Five labs

Use a cluster you can break: kind, k3d, or a sandbox namespace. Do not use production.

1. Run `check-client.sh`. If `context=set`, run `kubectl config current-context` yourself and decide whether that cluster is fair game.
2. `kubectl get ns` and `kubectl auth can-i get pods -n default`. Expected: `yes` or `no`. Either answer is the lesson.
3. Apply a Deployment of a public image you trust with `--dry-run=server` if can-i says you may create deployments. Read the output. Then apply for real only in the sandbox.
4. Point the container command at `exit 1`. Watch CrashLoopBackOff with `describe` and `logs --previous`. Fix the command. Do not delete the Deployment as the first move.
5. `kubectl rollout history` and `kubectl rollout undo` on that sandbox Deployment. Expected: the previous pod template returns. Confirm with `kubectl get pods`.

Progression: context and can-i, then get and describe, then diff and server dry-run, then rollout undo, and only then drain or delete.

## What this post does not cover

Cluster build (EKS version upgrades, Karpenter), service mesh, and policy engines. Runtime seccomp and AppArmor are in the [container runtime security](/blog/container-runtime-security-seccomp-apparmor-eks-fargate/) post. This page will go stale on flag details. `kubectl COMMAND --help` wins.

## What to do this week

1. Run the client check and write down the context you use for staging.
2. Practice `logs --previous` on a sandbox crash loop once, before a real order worker does it.
3. Remove any personal kubeconfig entry you no longer need. The file is a credential.
4. If deploys go through Git, read the [GitOps](/blog/aws-gitops-eks-argocd-flux-2026/) note and keep kubectl for diagnosis.

## Quick reference

| I need to | Command | Risk |
| --- | --- | --- |
| Know the cluster | `kubectl config current-context` | **Read-only** |
| See why a pod waits | `kubectl describe pod` | **Read-only** |
| See the crashed process | `kubectl logs --previous` | **Read-only** |
| Preview a write | `kubectl apply --dry-run=server` | Needs write RBAC. Should not persist. |
| Undo a Deployment | `kubectl rollout undo` | **Remote mutation** |
| Evict a node | `kubectl drain` | **Remote mutation** |

You should be able to say which cluster you are on, separate Pending from CrashLoopBackOff, and refuse cluster-admin as a debugging step.

## Further reading

- [kubectl reference](https://kubernetes.io/docs/reference/kubectl/)
- [kubectl quick reference](https://kubernetes.io/docs/reference/kubectl/quick-reference/)
- [Debug tasks](https://kubernetes.io/docs/tasks/debug/)
- [Dry-run authorization](https://kubernetes.io/docs/reference/using-api/api-concepts/#dry-run-authorization)
- Series: [Git](/blog/mastering-git-commands/), [Linux](/blog/mastering-linux-commands/), [AWS CLI](/blog/mastering-aws-cli/), [Docker](/blog/mastering-docker-commands/), [Bedrock CLIs](/blog/mastering-bedrock-cli/), [AI agent tools](/blog/mastering-ai-agent-tools/)

[Contact us](/contact-us/) or see [DevOps pipeline setup](/services/devops-pipeline-setup/) if production kubectl access is shared from one kubeconfig and nobody owns the RBAC review.

## FAQ

### When should you not grant cluster-admin to debug a pod?
As a routine fix. cluster-admin can read every secret and delete every namespace. Use kubectl auth can-i to see the missing verb, then ask for that verb on that resource in that namespace.


### Does kubectl diff only read the cluster?
kubectl diff submits a server-side dry run. The API server authorizes it like a write: patch on an existing object, create on a new one. A viewer role can fail diff even though nothing is persisted. Client-side dry-run (kubectl apply --dry-run=client) does not send that write.


### What is the difference between delete and force delete?
A normal delete asks the pod to shut down and respects the grace period. Force delete with a zero grace period removes the API object even if the kubelet has not confirmed the process is gone. Use it only when the node is dead and you accept that a process might still be running there.


### Helm and kubectl are the same tool?
No. kubectl talks to the Kubernetes API. Helm renders charts and stores release state. This page's authoring machine did not have helm on PATH. Treat Helm commands as a separate optional section and run helm --help before you rely on a flag.


### Does this page list every kubectl command?
No. The full tree is kubectl --help and the Kubernetes kubectl reference. This page is the operations set for workloads, network, auth, and recovery.


---

*Source: https://www.factualminds.com/blog/mastering-kubernetes-commands/*
