Bartosz Zaczyński’s Real Python roundup, updated October 5, is not a PyTorch release post. 2.14.0 landed September 2. The digest is what to notice if you upgraded without reading the notes. When the input to torch.clamp() or torch.clamp_min() sits exactly on a scalar bound, the gradient is now zero instead of one. In 2.13 the same snippet printed tensor(1.). That is not an exception. That is a quieter training run.
We already treated vLLM as 39 percent of the speed of light and verl 0.9.0 shipping a vLLM image. This week is pins: autograd, the vLLM runner default, and a Transformers parallel API that will not import the old way.
The clamp gradient is now zero on the bound
Real Python’s example is small enough to steal. x is torch.tensor(0.0, requires_grad=True). torch.clamp_min(x, 0.0).backward(). x.grad is tensor(0.). They suggest checking it in a throwaway env on Python 3.14 with uv run --python 3.14 --with torch==2.14.0 python. If you are still on 2.13, the last line is tensor(1.).
ReLU-shaped kinks always had a convention at zero. 2.14 picked the other convention for clamp at the bound. “Mathematically cleaner,” the digest says. Cleaner is not “your run matches last month.” If a training job is in a reproducibility gate, pin torch==2.13.x until you compare a few seeds, or pin 2.14 and accept the new zeros.
This will not show up as a shape error. It will show up as a slightly different curve on the steps where a lot of activations sit on the clip. If you clamp rewards, images, or logits as part of a backward path, you are in the set. If you only clamp for display, you are not.
Do not upgrade production in the same PR that “just bumps torch.” Read the 2.14 note for clamp, then bump.
Ties, fmin, and reproducibility
When the bound is itself a tensor that requires gradients, a tie now splits the gradient evenly between the input and the bound. fmin() and fmax() follow the same rule. That is a second silent change in the same paragraph. Elementwise min/max with two live tensors is common in custom losses. If you wrote a test that asserted an exact grad tensor, it can fail without a traceback from PyTorch, only from you.
Real Python’s advice is the boring one. Pin the version, or compare a few runs before you take 2.14. Exact bit-reproducibility across minor versions was already a hobby. This change makes the hobby visible.
If you need the old behavior, the digest does not offer a flag. I am not going to invent legacy_clamp_grad=True. Stay on 2.13 until you have a reason. If you need 2.14 for something else, write a tiny unit test around your clip and look at the grad once.
Transformers loading GGUF while llama.cpp owns the box was a serving split. This is a training split. Different files, same rule: the default moved.
vLLM 0.29 made V2 the default
Same digest, next sentences. vLLM 0.29.0 made Model Runner V2 the default and plans to remove V1 in 0.32. If your compose file still talks about V1 flags, that is a calendar. 0.32 is the deletion. 0.29 is the surprise if you pip install -U on a Monday.
We compared vLLM, TensorRT-LLM, and ONNX Runtime as engines. Defaults inside one engine are how those comparisons rot. A throughput number you printed against V1 is not a V2 number. Rerun the bench or keep the old extra.
I do not have the 0.29 changelog in this roundup beyond the default and the 0.32 removal plan. That is enough to act. Freeze vllm==0.28.x if you are in a bake-off. Read the V2 notes before you take 0.29 on a GPU that was stable.
verl’s October image with vLLM inside is a different pin. Do not assume that image is 0.29. Check the tag.
Transformers 5.16 replaced the parallel API
Hugging Face Transformers 5.16 replaced its legacy tensor-parallel API with a DTensor-based one. Real Python calls that a breaking change if you shard models yourself. If you only from_pretrained and generate, you may not care. If you have a custom device_map plus old TP helpers, you care.
5.17, in the same breath, added Tencent’s 780-billion-parameter Hy4 preview. Preview is not a reason to unpin 5.16. It is a reason to keep the extra isolated.
LangChain 1.4.0 added a native langchain.mcp namespace for Model Context Protocol, on top of August’s MCP rewrite. OpenAI’s Agents API entered public beta on September 10 as a hosted harness for sessions, context compaction, and recovery. Those are adjacent. They are not autograd. If your job this week is a training pin, do not also “just” bump LangChain.
The through-line in Zaczyński’s ML block is that the breakage did not always raise. PyTorch changed a number. vLLM changed a default. Transformers changed an API you only hit if you shard. Three different failure modes. One digest.
Pin the digest, not the vibes
Python 3.15.0 was due October 1 under PEP 790. Hugo van Kemenade shipped 3.15.0rc3 on October 2 instead, about 156 fixes from 82 contributors since rc2, because lazy-import blockers showed up late. Final is Friday, October 9. ABI frozen since August. That is interpreter news. It is in the same article so you do not mix it with torch 2.14. 3.15rc3 does not fix your clamp. torch 2.14 does not need 3.15.
If you bump three of these in one afternoon because a roundup listed them, you will not know which change moved the loss. One pin per PR. Paste the Real Python sentences into the PR body so the next person knows it was a silent grad, not a new kernel.
Conference calendar in that piece: PyBay October 3, PyCon Africa October 7 to 11, PyCon Greece October 12 to 13. None of that patches clamp. Skip it unless you are booking a flight.
Do not mix this pin with 3.15rc3
The same Real Python file spends its lede on Python 3.15.0rc3. Release manager Hugo van Kemenade delayed the final from October 1 to Friday, October 9, because lazy imports still had last-minute blockers. One bug: lazy import a.b as c treated b as an attribute of a instead of importing the submodule. Another: touching one lazily imported submodule also imported a sibling that should have stayed lazy. About 156 fixes from 82 contributors since rc2. Feature set frozen since May. ABI frozen since August.
That is not your clamp. If you are validating torch 2.14, stay on the Python you already train on. Mixing an rc interpreter with a new autograd convention is how you file a ghost bug. uv can run 3.15 when you mean to test 3.15. It can run 3.14 with torch 2.14 when you mean to test clamp. Use two commands.
SQLAlchemy 2.1.0 shipped September 24 in that digest too: greenlet no longer default, postgresql:// now psycopg 3, Python 3.11 minimum, tstring() on 3.14+. If your training box also talks to Postgres, that is a second ticket. Not this one.
LangChain’s langchain.mcp namespace and OpenAI’s Agents API beta are the same article’s gravity well. They will eat the PR if you let them. MCP is how agents attach tools. Clamp is how a tensor gets a grad. Different files.
If you serve with vLLM and train with PyTorch on one clone, freeze them on different branches or different extras. 0.29’s V2 default will change serving traces. 2.14’s clamp will change training traces. A joint bump is unreadable.
What to run this week
Throwaway check, from the digest:
uv run --python 3.14 --with torch==2.14.0 python
Then the three lines: tensor 0.0 with requires_grad, clamp_min against 0.0, backward, print x.grad. Confirm you see tensor(0.). On 2.13, confirm tensor(1.). Put that in a comment next to any clip in your loss.
pip freeze | grep -E 'torch==|vllm==|transformers==' and write the three versions in the ticket. If vLLM is already 0.29, schedule a V2 read before 0.32. If Transformers is 5.16+ and you shard, grep your repo for the old TP symbols before you roll more GPUs.
Do not take Hy4 preview, Agents API beta, and a clamp change in the same window. The digest is a menu. Order one plate.
One digest, three tickets
Print three tickets from the October 5 file and refuse to close them together. Ticket one: torch 2.14 clamp and fmin/fmax ties. Ticket two: vLLM 0.29 V2 default and the 0.32 V1 removal. Ticket three: Transformers 5.16 DTensor TP. If a fourth ticket appears because someone saw Hy4 or Agents API in Slack, it waits.
The roundup also notes PyCon Ireland moved from October 17 to November 21 in Dublin after the venue fell through, and PyCon Estonia canceled 2026. That is travel. It does not belong on the freeze file.
When you paste Real Python into a PR, paste the clamp example, not the table of contents. Reviewers will skim. Give them the tensor.
If a unit test asserts x.grad == 1 on a clipped zero, it is now a 2.13 test. Rename it or change the assertion and leave a comment with the Real Python date. Future you will not remember why the number flipped.
Hy4 at 780 billion in 5.17 is a preview line in a digest. It is not a reason to unpin serving. Agents API beta on September 10 is a hosted harness. It does not change clamp_min. Keep those tabs closed until the torch pin merges.
PyData Global is online December 8 to 10 in the same calendar block. Still travel-or-Zoom, still not a pin. If your team wants a reading group, read the 2.14 clamp note, not a keynote abstract.
If serving and training share a lockfile, split it. One extra for train, one for serve. The digest listed them in one paragraph because they happened in one month, not because they belong in one install.
Keep the three tickets on a sticky note on the monitor until they merge. A roundup tab is not a version pin.
2.14 will be fine for most people. The people it will not be fine for are the ones who needed the old one at the bound and never wrote it down. Write it down now.
Discussion
Leave a comment
No comments yet
Be the first to start the conversation.