<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hardware on kenji.blog</title><link>http://kenji.blog/en/categories/hardware/</link><description>Recent content in Hardware on kenji.blog</description><generator>Hugo -- gohugo.io</generator><language>en</language><copyright>kenjinote</copyright><lastBuildDate>Sat, 12 Sep 2026 12:00:00 +0900</lastBuildDate><atom:link href="http://kenji.blog/en/categories/hardware/index.xml" rel="self" type="application/rss+xml"/><item><title>For Long Coding Sessions! 5 Recommended Mechanical Keyboards for Engineers</title><link>http://kenji.blog/en/p/engineer-mechanical-keyboard-recommendations/</link><pubDate>Sat, 12 Sep 2026 12:00:00 +0900</pubDate><guid>http://kenji.blog/en/p/engineer-mechanical-keyboard-recommendations/</guid><description>&lt;img src="http://kenji.blog/p/engineer-mechanical-keyboard-recommendations/img/eyecatch.jpg" alt="Featured image of post For Long Coding Sessions! 5 Recommended Mechanical Keyboards for Engineers" />&lt;h1 id="for-long-coding-sessions-5-recommended-mechanical-keyboards-for-engineers">For Long Coding Sessions! 5 Recommended Mechanical Keyboards for Engineers
&lt;/h1>&lt;p>For professionals working in the IT industry, such as programmers, system engineers, and data scientists, a keyboard is not merely an input device. It is an &amp;ldquo;interface for converting thoughts into code,&amp;rdquo; and the most important work tool that they directly touch for hours every day.&lt;/p>
&lt;p>Continuing to use a poor quality keyboard not only causes a decrease in typing speed, but also increases excessive strain on the wrists and finger joints, and consequently, the risk of tendonitis (such as carpal tunnel syndrome). Conversely, obtaining a highly customizable keyboard that fits well in your hand and has a good typing feel is the &amp;ldquo;best investment&amp;rdquo; that greatly improves both productivity and health.&lt;/p>
&lt;p>In this article, aimed at engineers, we will go beyond mere &amp;ldquo;recommendations&amp;rdquo; and thoroughly explain everything from the physics of keyboards to internal electronic circuits, and the latest firmware technology. Based on that, we will introduce 5 ultimate keyboards that truly withstand practical use.&lt;/p>
&lt;h2 id="1-physics-and-mechanisms-of-key-switches">1. Physics and Mechanisms of Key Switches
&lt;/h2>&lt;p>The most important element that determines the typing feel of a keyboard is the &amp;ldquo;key switch&amp;rdquo;. Mechanical keyboard switches consist of a spring and a contact mechanism, and their physical characteristics are transmitted to our fingertips as feedback.&lt;/p>
&lt;h3 id="11-hookes-law-and-spring-constant">1.1 Hooke&amp;rsquo;s Law and Spring Constant
&lt;/h3>&lt;p>The actuation force of a mechanical switch is mainly determined by the characteristics of the spring installed inside. The behavior of this spring can be approximately represented by &amp;ldquo;Hooke&amp;rsquo;s Law&amp;rdquo; in classical mechanics.&lt;/p>
$$ F = -k x $$&lt;p>Here, $F$ is the restoring force (the repulsive force felt by the finger), $k$ is the spring constant, and $x$ is the pushed distance (stroke).
In the case of linear switches (such as red or black switches), they follow this Hooke&amp;rsquo;s Law almost faithfully, having a linear characteristic where the repulsive force increases proportionally the more you push down.&lt;/p>
&lt;h3 id="12-integral-calculation-of-actuation-energy">1.2 Integral Calculation of Actuation Energy
&lt;/h3>&lt;p>The point at which a key is recognized as &amp;ldquo;input&amp;rdquo; is called the Actuation Point. The energy (amount of work) $E$ spent by a finger from the start of pushing the key until reaching the actuation point $x_a$ is expressed by the integral of force over distance.&lt;/p>
$$ E = \int_{0}^{x_a} F(x) \, dx $$&lt;p>In the case of tactile switches (brown switches) or clicky switches (blue switches), there is physical resistance where the contacts rub against each other (tactile bump), so $F(x)$ is not a simple linear function, but a function that peaks non-linearly at a specific stroke position.&lt;/p>
&lt;pre class="mermaid">
flowchart TD
A[&amp;#34;Start of finger press&amp;#34;] --&amp;gt; B{&amp;#34;Switch type&amp;#34;}
B --&amp;gt;|Linear| C[&amp;#34;Resistance increases linearly&amp;#34;]
B --&amp;gt;|Tactile| D[&amp;#34;Physical resistance (bump) in the middle&amp;#34;]
B --&amp;gt;|Clicky| E[&amp;#34;Sound generation mechanism operates simultaneously with the bump&amp;#34;]
C --&amp;gt; F[&amp;#34;Reach Actuation Point&amp;#34;]
D --&amp;gt; F
E --&amp;gt; F
F --&amp;gt; G[&amp;#34;Bottom Out&amp;#34;]
&lt;/pre>
&lt;p>When an engineer codes for a long time, if this $E$ (actuation energy) is too large, the fingers get tired easily, and if it is too small, mistypes (accidental hits) increase. Generally, a switch with an actuation force of about 45g to 55g is considered to have a good balance of fatigue reduction and accuracy, and is preferred by many engineers.&lt;/p>
&lt;h3 id="13-state-of-the-art-switch-technology-electrostatic-capacitive-non-contact-and-hall-effect">1.3 State-of-the-Art Switch Technology: Electrostatic Capacitive Non-Contact and Hall Effect
&lt;/h3>&lt;p>There are also more advanced switch technologies that do not have physical metal contacts.&lt;/p>
&lt;p>&lt;strong>Electrostatic Capacitive Non-Contact (Topre)&lt;/strong>
Using a conical spring and a rubber dome, it determines input by detecting the change in electrostatic capacity caused by pushing down. Because there is no physical contact, wear is extremely low, and chattering (the phenomenon of multiple inputs being registered with a single press) does not occur. The unique &amp;ldquo;thock&amp;rdquo; typing feel provided by the rubber dome has a charm that you cannot leave once you experience it.&lt;/p>
&lt;p>&lt;strong>Magnetic Switch (Hall Effect)&lt;/strong>
Using the Hall effect, it reads the change in magnetic flux density as voltage when a magnet embedded in the stem (axis) approaches a Hall sensor on the circuit board.
The electromotive force $V_H$ due to the Hall effect is expressed by the following equation.&lt;/p>
$$ V_H = R_H \left( \frac{I \cdot B}{t} \right) $$&lt;p>Here, $R_H$ is the Hall coefficient, $I$ is the current, $B$ is the magnetic flux density, and $t$ is the thickness of the conductor. With this technology, the depth of the keystroke can be continuously obtained as an analog value, enabling incredible control such as &amp;ldquo;Actuation Point Adjustment&amp;rdquo; (changing the actuation point in units of 0.1mm) and &amp;ldquo;Rapid Trigger&amp;rdquo; (turning off the key the moment you start to release it).&lt;/p>
&lt;h2 id="2-keyboard-electronic-circuits-and-performance-metrics">2. Keyboard Electronic Circuits and Performance Metrics
&lt;/h2>&lt;p>Even if the switches are excellent, if the performance of the electronic circuits and microcontroller (MCU) that process them is low, the best performance cannot be demonstrated.&lt;/p>
&lt;h3 id="21-matrix-scanning-and-polling-rate">2.1 Matrix Scanning and Polling Rate
&lt;/h3>&lt;p>Inside a keyboard, there are anywhere from dozens to over 100 switches, but since the number of pins on a microcontroller is limited, it is impossible to connect every switch to an individual pin. Therefore, switches are wired in a grid (matrix) of Rows and Columns, and by scanning them at high speed, it determines which key was pressed.&lt;/p>
&lt;pre class="mermaid">
flowchart LR
M[&amp;#34;Microcontroller (MCU)&amp;#34;] --&amp;gt;|Switch Row output to High/Low| R1[&amp;#34;Row 1&amp;#34;]
M --&amp;gt; R2[&amp;#34;Row 2&amp;#34;]
R1 --&amp;gt; S11[&amp;#34;Switch 1,1&amp;#34;] &amp;amp; S12[&amp;#34;Switch 1,2&amp;#34;]
R2 --&amp;gt; S21[&amp;#34;Switch 2,1&amp;#34;] &amp;amp; S22[&amp;#34;Switch 2,2&amp;#34;]
S11 &amp;amp; S21 --&amp;gt; C1[&amp;#34;Column 1&amp;#34;]
S12 &amp;amp; S22 --&amp;gt; C2[&amp;#34;Column 2&amp;#34;]
C1 &amp;amp; C2 --&amp;gt;|Detect and read voltage| M
&lt;/pre>
&lt;p>&lt;strong>Polling Rate&lt;/strong> is the frequency at which the keyboard reports its &amp;ldquo;current key status&amp;rdquo; to the PC. Standard keyboards are 125Hz (once every 8ms), but high-end models have ultra-high-speed communication such as 1000Hz (once every 1ms) or recently 8000Hz (once every 0.125ms).
For coding, 1000Hz is more than enough performance, but it leads to a sense of security that prevents missed keystrokes during ultra-high-speed typing.&lt;/p>
&lt;h3 id="22-n-key-rollover-nkro-and-anti-ghosting">2.2 N-Key Rollover (NKRO) and Anti-Ghosting
&lt;/h3>&lt;p>&lt;strong>N-Key Rollover (NKRO)&lt;/strong> is a feature where, when multiple keys are pressed simultaneously, all of them are accurately recognized. In the past, due to USB connection restrictions, there were limits like &amp;ldquo;up to 6 keys,&amp;rdquo; but modern high-end keyboards have achieved virtually unlimited simultaneous presses (Full NKRO) by cleverly utilizing USB HID reports.&lt;/p>
&lt;p>For engineers who heavily use complex shortcuts in editors like Vim or Emacs (e.g., &lt;code>Ctrl + Shift + Alt + any key&lt;/code>), a complete NKRO is a prerequisite.&lt;/p>
&lt;h3 id="23-debounce-delay">2.3 Debounce Delay
&lt;/h3>&lt;p>Mechanical switches with metal contacts experience a &amp;ldquo;bounce phenomenon&amp;rdquo; where the contacts slightly bounce when pressed or released. The processing time for the microcontroller to ignore this is the &lt;strong>Debounce Delay&lt;/strong>. Normally, an intentional delay of about 5ms to 20ms is provided, but in the aforementioned electrostatic capacitive non-contact systems and magnetic switches, since there is no physical contact noise, the debounce delay can be set to zero (or minimal), achieving overwhelming response.&lt;/p>
&lt;h2 id="3-firmware-and-customizability-qmk--via">3. Firmware and Customizability (QMK / VIA)
&lt;/h2>&lt;p>If the hardware is the &amp;ldquo;body&amp;rdquo;, the firmware is the &amp;ldquo;brain&amp;rdquo; of the keyboard. Modern high-end keyboards for engineers have the ability not just to send keycodes, but to execute advanced programs.&lt;/p>
&lt;h3 id="31-qmk-firmware">3.1 QMK Firmware
&lt;/h3>&lt;p>&lt;strong>QMK (Quantum Mechanical Keyboard)&lt;/strong> is an open-source keyboard firmware. Written in C, it literally allows you to do &amp;ldquo;anything,&amp;rdquo; from changing keymaps to creating macros and controlling LED animations.&lt;/p>
&lt;h3 id="32-advanced-key-assignment-features">3.2 Advanced Key Assignment Features
&lt;/h3>&lt;p>Among the features provided by QMK, the following in particular explosively increase engineer productivity.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Layers:&lt;/strong> Just like switching between &amp;ldquo;letters&amp;rdquo; and &amp;ldquo;numbers&amp;rdquo; on a smartphone keyboard, the entire layout of the keyboard is switched to a different one only while a specific key (such as the Fn key) is pressed. This makes it possible to input arrow keys, macros, and symbols without moving your hands from the home row.&lt;/li>
&lt;li>&lt;strong>Mod-Tap:&lt;/strong> Assigns different roles to a single key for &amp;ldquo;when tapped briefly&amp;rdquo; and &amp;ldquo;when held down.&amp;rdquo; For example, setting the space bar to &amp;ldquo;Space on tap, Shift on hold&amp;rdquo; (Space Cadet Shift) enables effective use of the thumbs.&lt;/li>
&lt;li>&lt;strong>Home Row Mods:&lt;/strong> A technique of assigning modifiers (Ctrl, Shift, Alt, GUI) on hold to keys on the home row (ASDF, JKL;, etc.). This eliminates the need to overwork your pinky to reach the Ctrl key, dramatically reducing wrist fatigue for Vim and Emacs users.&lt;/li>
&lt;/ul>
&lt;h3 id="33-real-time-configuration-with-via--vial">3.3 Real-time Configuration with VIA / VIAL
&lt;/h3>&lt;p>The drawback of QMK was that &amp;ldquo;every time you change settings, you have to compile the source code and flash (write) the firmware.&amp;rdquo; &lt;strong>VIA&lt;/strong> and &lt;strong>VIAL&lt;/strong> solved this. These allow you to access the keyboard from a GUI application (or a web browser) and rewrite the keymap in real-time without rebooting.&lt;/p>
&lt;h2 id="4-ergonomics-and-the-science-of-layouts">4. Ergonomics and the Science of Layouts
&lt;/h2>&lt;p>The typical &amp;ldquo;row-staggered (keys are staggered by row)&amp;rdquo; layout is a remnant to prevent the physical arms of typewriters from tangling, and is not based on the structure of the human hand.&lt;/p>
&lt;pre class="mermaid">
pie title Engineer&amp;#39;s Ideal Keyboard Layout Preferences (Estimated Data)
&amp;#34;Row Staggered (Conventional)&amp;#34; : 45
&amp;#34;Alice Layout (Ergonomic)&amp;#34; : 15
&amp;#34;Ortholinear (Grid Layout)&amp;#34; : 10
&amp;#34;Column Staggered (Split)&amp;#34; : 30
&lt;/pre>
&lt;p>There are layouts that are more ergonomically considerate, such as the following.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Ortholinear:&lt;/strong> A layout where keys are arranged in a perfectly straight grid vertically and horizontally. The bending and stretching of fingers becomes linear, reducing wasted finger movement.&lt;/li>
&lt;li>&lt;strong>Columnar Stagger:&lt;/strong> A layout where vertical columns are staggered according to the length of human fingers (middle finger is long, pinky is short). You can type with a natural hand shape.&lt;/li>
&lt;li>&lt;strong>Split:&lt;/strong> Since the left and right hands can be placed completely apart, you can type in a natural posture with your chest open and shoulders relaxed, demonstrating immense effectiveness in preventing stiff shoulders and straight neck.&lt;/li>
&lt;/ul>
&lt;h2 id="5-5-ultimate-mechanical-keyboards-recommended-for-engineers">5. 5 Ultimate Mechanical Keyboards Recommended for Engineers
&lt;/h2>&lt;p>Based on physics, electronic circuits, firmware, and ergonomics, we have carefully selected 5 keyboards for true professionals that can withstand long hours of coding.&lt;/p>
&lt;hr>
&lt;h3 id="1-keychron-q-series-q1-pro--q8-etc---the-gateway-to-the-world-of-custom-keyboards">1. Keychron Q Series (Q1 Pro / Q8, etc.) - The Gateway to the World of Custom Keyboards
&lt;/h3>&lt;p>Keychron, originating from Hong Kong, is driving the recent custom keyboard boom. Among them, the &amp;ldquo;Q Series&amp;rdquo; features a heavy full aluminum body and a &amp;ldquo;Gasket Mount&amp;rdquo; structure that tunes the typing sound to the utmost limit.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Switches:&lt;/strong> Mechanical (Hot-swappable. Switches can be freely exchanged)&lt;/li>
&lt;li>&lt;strong>Firmware:&lt;/strong> Fully compatible with QMK/VIA&lt;/li>
&lt;li>&lt;strong>Features:&lt;/strong> A toggle switch for both macOS and Windows compatibility. You can choose your preferred layout, such as the Q8 with an Alice layout or the Q1 with a 75% layout.&lt;/li>
&lt;li>&lt;strong>Benefits for Engineers:&lt;/strong> Despite being a pre-built product, you can immediately experience superb typing feel and customizability comparable to a custom-built keyboard right out of the box. It is ideal for setting up a Vim-like arrow layer using VIA.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h3 id="2-hhkb-studio---the-all-in-one-pointing-device-for-hackers">2. HHKB Studio - The All-in-One Pointing Device for Hackers
&lt;/h3>&lt;p>The &amp;ldquo;Happy Hacking Keyboard (HHKB)&amp;rdquo; is a legendary keyboard born for UNIX programmers. The latest &amp;ldquo;HHKB Studio&amp;rdquo; has evolved further by adopting specially developed silent mechanical switches instead of the conventional electrostatic capacitive non-contact method.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Switches:&lt;/strong> Linear silent mechanical switches (manufactured by Kailh, hot-swappable)&lt;/li>
&lt;li>&lt;strong>Features:&lt;/strong> A pointing stick (TrackPoint) in the center of the keyboard, 4 gesture pads.&lt;/li>
&lt;li>&lt;strong>Benefits for Engineers:&lt;/strong> You can complete mouse cursor operations, scrolling, and window switching without ever taking your hands off the home row. Once you experience this &amp;ldquo;everything is completed at your fingertips&amp;rdquo; experience, you can never go back to reaching for a mouse with your right hand.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h3 id="3-zsa-moonlander--ergodox-ez---ultimate-split-ergonomics">3. ZSA Moonlander / ErgoDox EZ - Ultimate Split Ergonomics
&lt;/h3>&lt;p>The pinnacle of split keyboards developed by Canada&amp;rsquo;s ZSA. The left and right sides are independent and can be placed according to your shoulder width, surprisingly reducing the burden on your shoulders and neck even during long hours of typing.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Switches:&lt;/strong> Mechanical (Cherry MX compatible, hot-swappable)&lt;/li>
&lt;li>&lt;strong>Firmware:&lt;/strong> QMK based (using its own powerful GUI tool &amp;ldquo;Oryx&amp;rdquo;)&lt;/li>
&lt;li>&lt;strong>Features:&lt;/strong> Columnar staggered layout, dedicated thumb cluster keys, and legs for tenting (tilting) come standard.&lt;/li>
&lt;li>&lt;strong>Benefits for Engineers:&lt;/strong> By assigning Enter, Space, Backspace, and Layer switching to your thumbs, the burden on the weakest pinky fingers is drastically reduced. It is a savior device for engineers suffering from carpal tunnel syndrome.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h3 id="4-realforce-r3---domestic-reliability-and-supreme-typing-feel-electrostatic-capacitive-non-contact">4. REALFORCE R3 - Domestic Reliability and Supreme Typing Feel (Electrostatic Capacitive Non-Contact)
&lt;/h3>&lt;p>A Japanese masterpiece boasted by Topre. The track record of being used for many years in professional fields such as financial institutions is not just for show. From the R3 generation, it also supports Bluetooth connection.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Switches:&lt;/strong> Electrostatic Capacitive Non-Contact (Topre)&lt;/li>
&lt;li>&lt;strong>Features:&lt;/strong> With the APC (Actuation Point Changer) function, the actuation point can be set per key from 0.8mm, 1.5mm, 2.2mm, and 3.0mm.&lt;/li>
&lt;li>&lt;strong>Benefits for Engineers:&lt;/strong> The smooth key touch due to the absence of physical contacts is called &amp;ldquo;feather touch&amp;rdquo;, and the repulsive stress on the fingers is kept to a minimum even during long coding sessions. It is possible to customize it so that only keys pressed by the pinky (like A or Enter) have a shallow actuation point (0.8mm), allowing them to react with just a light touch.&lt;/li>
&lt;/ul>
&lt;hr>
&lt;h3 id="5-wooting-60he---revolutionary-response-brought-by-magnetic-switches">5. Wooting 60HE - Revolutionary Response Brought by Magnetic Switches
&lt;/h3>&lt;p>Originally developed for e-sports gamers, its innovative technology is also highly evaluated by engineers seeking the fastest typing and response.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Switches:&lt;/strong> Lekker Switch (Hall effect magnetic switch)&lt;/li>
&lt;li>&lt;strong>Features:&lt;/strong> Rapid trigger function, actuation point adjustable in 0.1mm increments from 0.1mm to 4.0mm.&lt;/li>
&lt;li>&lt;strong>Benefits for Engineers:&lt;/strong> Utilizing analog input, eccentric settings (Dynamic Keystroke) like &amp;ldquo;lowercase if pushed slightly, uppercase if pushed deeply (in combination with Shift)&amp;rdquo; are possible. In addition, since the key turns off the moment the finger is lifted even slightly, it prevents unintended continuous key inputs during high-speed typing, providing an unparalleled accurate input experience.&lt;/li>
&lt;/ul>
&lt;h2 id="conclusion">Conclusion
&lt;/h2>&lt;p>Choosing a keyboard is a process of &amp;ldquo;optimizing your own interface&amp;rdquo; throughout your career as an engineer. From the feel of physical springs that obey Hooke&amp;rsquo;s Law, to actuation energy calculated by integration, macro building by QMK, and ultimate ergonomics, the depth to be pursued is bottomless.&lt;/p>
&lt;p>The 5 keyboards introduced this time (Keychron, HHKB Studio, Moonlander, REALFORCE, Wooting) are all masterpieces aiming for the &amp;ldquo;best input experience&amp;rdquo; with different approaches. By all means, please find your best partner according to your typing style and the physical troubles you have.&lt;/p>
&lt;p>An investment in a keyboard will surely bring returns to you as &amp;ldquo;millions of lines of bug-free code&amp;rdquo;.&lt;/p></description></item><item><title>Gadgets and Monitor Settings to Reduce Programmer's Eye Strain</title><link>http://kenji.blog/en/p/programmer-eye-strain-relief/</link><pubDate>Sat, 12 Sep 2026 12:00:00 +0900</pubDate><guid>http://kenji.blog/en/p/programmer-eye-strain-relief/</guid><description>&lt;img src="http://kenji.blog/p/programmer-eye-strain-relief/img/eyecatch.jpg" alt="Featured image of post Gadgets and Monitor Settings to Reduce Programmer's Eye Strain" />&lt;p>For programmers and software engineers, the &amp;ldquo;eyes&amp;rdquo; are their most important and heavily abused tools of the trade. Spending 8 to 10 hours a day, sometimes even more, constantly looking at editors, terminals, and browser screens, almost every engineer faces &amp;ldquo;Computer Vision Syndrome&amp;rdquo; (CVS) or eye strain.&lt;/p>
&lt;p>Generally, countermeasures against eye strain tend to end with superficial advice such as &amp;ldquo;use eye drops,&amp;rdquo; &amp;ldquo;take appropriate breaks,&amp;rdquo; or &amp;ldquo;wear blue light blocking glasses.&amp;rdquo; However, as engineers, we should identify the root cause of the problem and optimize it from the system (environment) layer.&lt;/p>
&lt;p>In this article, we will thoroughly dissect the mechanics of a programmer&amp;rsquo;s eye strain from the perspectives of physics (optics), biochemistry, ergonomics, and display hardware architecture. We will delve deeply into the ultimate monitor settings and gadgets to relieve it, using mathematical formulas and illustrations.&lt;/p>
&lt;hr>
&lt;h1 id="chapter-1-unraveling-the-mechanics-of-eye-strain-cvs-through-physics-and-biochemistry">Chapter 1: Unraveling the Mechanics of Eye Strain (CVS) Through Physics and Biochemistry
&lt;/h1>&lt;p>Computer Vision Syndrome (CVS) is not caused by a single factor. As shown in the pie chart below, various elements are complexly intertwined, leading to eye fatigue, pain, dry eyes, and overall bodily fatigue.&lt;/p>
&lt;pre class="mermaid">
pie title Causes of Computer Vision Syndrome (CVS)
&amp;#34;Blue Light &amp;amp; Glare&amp;#34; : 30
&amp;#34;Screen Flickering (PWM)&amp;#34; : 25
&amp;#34;Improper Contrast &amp;amp; Lighting&amp;#34; : 20
&amp;#34;Focus Fatigue (Ciliary Muscle)&amp;#34; : 15
&amp;#34;Dry Eyes (Reduced Blinking)&amp;#34; : 10
&lt;/pre>
&lt;p>Here, we will explain the &amp;ldquo;physical properties of light&amp;rdquo; and the &amp;ldquo;focus adjustment function of the eyeball,&amp;rdquo; which have particularly significant impacts.&lt;/p>
&lt;h2 id="11-physical-properties-of-blue-light-and-photon-energy">1.1 Physical Properties of Blue Light and Photon Energy
&lt;/h2>&lt;p>Blue light emitted from displays is located roughly in the wavelength band of $400 \text{ nm} \sim 490 \text{ nm}$. The reason why this puts a strain on the eyes can be explained by the &amp;ldquo;Planck-Einstein relation,&amp;rdquo; which is the foundation of quantum mechanics.&lt;/p>
&lt;p>The energy of light $E$ is expressed by the following formula:&lt;/p>
$$ E = h\nu = \frac{hc}{\lambda} $$&lt;p>Here, each variable has the following meaning:&lt;/p>
&lt;ul>
&lt;li>$E$ : Energy per photon (Joule)&lt;/li>
&lt;li>$h$ : Planck constant ($6.626 \times 10^{-34} \text{ J}\cdot\text{s}$)&lt;/li>
&lt;li>$c$ : Speed of light in a vacuum ($3.0 \times 10^8 \text{ m/s}$)&lt;/li>
&lt;li>$\lambda$ : Wavelength of light (m)&lt;/li>
&lt;li>$\nu$ : Frequency of light (Hz)&lt;/li>
&lt;/ul>
&lt;p>The important fact indicated by this formula is that &lt;strong>&amp;ldquo;the energy of light $E$ is inversely proportional to the wavelength $\lambda$.&amp;rdquo;&lt;/strong> In other words, blue light, which has the shortest wavelength among visible light, possesses extremely high energy. These high-energy photons are less likely to be absorbed or attenuated by the cornea or crystalline lens, reaching deep into the retina and applying strong oxidative stress to the photoreceptor cells.&lt;/p>
&lt;h2 id="12-chromatic-aberration-and-focus-shift">1.2 Chromatic Aberration and Focus Shift
&lt;/h2>&lt;p>Furthermore, from an optical perspective, differences in the wavelength of light create differences in the &amp;ldquo;refractive index.&amp;rdquo; The refractive index $n$ of a medium (such as the crystalline lens in this case) depends on the wavelength $\lambda$, and is approximated by Cauchy&amp;rsquo;s equation:&lt;/p>
$$ n(\lambda) = B + \frac{C}{\lambda^2} $$&lt;p>($B, C$ are constants unique to the medium)&lt;/p>
&lt;p>As can be seen from this formula, the shorter the wavelength $\lambda$ of the blue light, the larger the refractive index $n$. Therefore, even if red light is perfectly focused on the retina, blue light is strongly refracted and comes to a focus &lt;strong>in front of the retina&lt;/strong>.
When the brain recognizes this &amp;ldquo;image blurring due to blue light (chromatic aberration),&amp;rdquo; it constantly sends commands to the ciliary muscle to unconsciously try to refocus. This is a major factor in the unconscious fatigue of the eye muscles.&lt;/p>
&lt;h2 id="13-focus-adjustment-muscle-ciliary-muscle-and-thin-lens-equation">1.3 Focus Adjustment Muscle (Ciliary Muscle) and Thin Lens Equation
&lt;/h2>&lt;p>When we focus on fine text on a monitor, we adjust the thickness of the crystalline lens inside our eyes. The thin lens equation is as follows:&lt;/p>
$$ \frac{1}{f} = \frac{1}{a} + \frac{1}{b} $$&lt;ul>
&lt;li>$f$: Focal length of the crystalline lens&lt;/li>
&lt;li>$a$: Distance from the eye to the monitor (object distance)&lt;/li>
&lt;li>$b$: Distance from the crystalline lens to the retina (image distance: constant at about $24 \text{ mm}$ in an adult eyeball)&lt;/li>
&lt;/ul>
&lt;p>During programming, if the distance $a$ to the monitor is kept short (e.g., $40 \text{ cm} \sim 50 \text{ cm}$) for a long time, the focal length $f$ must be kept extremely short in order to form an accurate image on the retina (keeping $b$ constant). If the ciliary muscle remains extremely contracted for hours, the muscle falls into a state of spasm, causing severe eye strain accompanied by stiff shoulders and headaches.&lt;/p>
&lt;hr>
&lt;h1 id="chapter-2-hardware-display-selection-and-elimination-of-fatigue-factors">Chapter 2: Hardware Display Selection and Elimination of Fatigue Factors
&lt;/h1>&lt;p>To alleviate eye fatigue, it is necessary to first verify and improve the hardware specifications before tweaking software settings. In particular, the &amp;ldquo;dimming method&amp;rdquo; and &amp;ldquo;refresh rate&amp;rdquo; are points where you should not compromise.&lt;/p>
&lt;h2 id="21-the-terror-of-pwm-dimming-uncovering-invisible-flicker">2.1 The Terror of PWM Dimming: Uncovering Invisible Flicker
&lt;/h2>&lt;p>The technologies for adjusting the brightness of LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode) monitors are broadly divided into &amp;ldquo;DC (Direct Current) dimming&amp;rdquo; and &amp;ldquo;PWM (Pulse-Width Modulation) dimming.&amp;rdquo;&lt;/p>
&lt;p>PWM dimming is a technology that blinks the backlight LEDs at a high speed invisible to the human eye, artificially adjusting screen brightness through the ratio of its &amp;ldquo;on time&amp;rdquo; to &amp;ldquo;off time.&amp;rdquo; The average brightness $L$ determined by the duty cycle of the PWM is expressed by the following equation:&lt;/p>
$$ L = L_{max} \times \frac{T_{on}}{T_{on} + T_{off}} \times 100 \ (\%) $$&lt;ul>
&lt;li>$T_{on}$ : Time the LED is on&lt;/li>
&lt;li>$T_{off}$ : Time the LED is off&lt;/li>
&lt;li>$L_{max}$ : Maximum peak brightness&lt;/li>
&lt;/ul>
&lt;p>When the frequency of PWM dimming is low (e.g., $200 \text{ Hz} \sim 300 \text{ Hz}$), even if you do not consciously perceive the screen flickering, the brain and pupils unconsciously react to the blinking light, repeatedly dilating and constricting. This induces extreme fatigue, headaches, and even nausea.&lt;/p>
&lt;p>&lt;strong>[How to Detect PWM and Countermeasures]&lt;/strong>
To check if your monitor uses PWM dimming, launch the camera app on your smartphone, set it to &amp;ldquo;slow-motion video&amp;rdquo; mode, and record a white screen on your monitor (like a blank browser page). If dark horizontal bands (banding) appear moving across the video, that monitor employs low-frequency PWM dimming.
When programmers choose a monitor, they should absolutely select one where the specifications clearly state &lt;strong>&amp;ldquo;Flicker-Free (DC dimming).&amp;rdquo;&lt;/strong>&lt;/p>
&lt;h2 id="22-ophthalmological-impact-of-refresh-rate-hz-and-motion-blur">2.2 Ophthalmological Impact of Refresh Rate (Hz) and Motion Blur
&lt;/h2>&lt;p>The refresh rate is a numerical value (Hz) indicating how many times the monitor redraws the screen per second.
Standard office monitors are $60 \text{ Hz}$, but high refresh rate monitors like $120 \text{ Hz}$ or $144 \text{ Hz}$ have become popular in recent years. This is extremely beneficial not only for gamers but also for programmers.&lt;/p>
&lt;p>When scrolling through massive amounts of code or when a large volume of logs flows in the terminal, a $60 \text{ Hz}$ display will experience &amp;ldquo;motion blur&amp;rdquo; (afterimages) combined with the limitations of pixel response times. The eye unconsciously tries to capture the shape of the text and keep it in focus even during scrolling, but if the characters are blurred, the processing load on the brain&amp;rsquo;s visual cortex spikes dramatically.
With a display of $120 \text{ Hz}$ or higher, text remains clearly visible even while scrolling, significantly reducing the burden of these unconscious eye movements and focus adjustments.&lt;/p>
&lt;h2 id="23-panel-types-and-contrast-ratio-ips-va-oled">2.3 Panel Types and Contrast Ratio (IPS, VA, OLED)
&lt;/h2>&lt;p>The contrast ratio of a screen directly affects text legibility.
The &amp;ldquo;Weber-Fechner Law,&amp;rdquo; which states that the magnitude of human sensation is proportional to the logarithm of the stimulus, is expressed by the following formula:&lt;/p>
$$ p = k \ln \left( \frac{S}{S_0} \right) $$&lt;p>($p$: magnitude of sensation, $S$: physical magnitude of stimulus, $S_0$: threshold, $k$: constant)&lt;/p>
&lt;p>In other words, the human eye reacts more strongly to the &amp;ldquo;relative brightness ratio (contrast)&amp;rdquo; than to absolute brightness.
When reading syntax-highlighted code for long periods, VA panels ($3000:1$) with deep blacks (high contrast ratio) or OLED panels ($1,000,000:1$ and up) that can completely turn off individual pixels make character outlines very clear and improve legibility.
However, as explained later, looking at an extremely high-contrast screen in a pitch-black room causes the pupils to constrict too much, leading to fatigue instead, so a balance with ambient light is essential.&lt;/p>
&lt;p>The chart below compares the conceptual emission spectrum of a standard LCD monitor with modern OLED (low blue light design).&lt;/p>
&lt;pre class="mermaid">
xychart-beta
title Blue Light Emission Spectrum Comparison
x-axis &amp;#34;Wavelength (nm)&amp;#34; [400, 420, 440, 460, 480, 500]
y-axis &amp;#34;Relative Intensity&amp;#34; 0 --&amp;gt; 100
bar &amp;#34;Standard LCD (W-LED)&amp;#34; [10, 30, 95, 80, 40, 20]
line &amp;#34;Modern OLED / Low Blue Light&amp;#34; [5, 10, 40, 75, 55, 30]
&lt;/pre>
&lt;hr>
&lt;h1 id="chapter-3-monitor-calibration-and-os--software-settings">Chapter 3: Monitor Calibration and OS / Software Settings
&lt;/h1>&lt;p>Equally as important as hardware selection is color space management and calibration on the OS side.&lt;/p>
&lt;h2 id="31-the-trap-of-color-gamut-srgb-vs-dci-p3-and-icc-profiles">3.1 The Trap of Color Gamut (sRGB vs DCI-P3) and ICC Profiles
&lt;/h2>&lt;p>Modern monitors often boast a &amp;ldquo;wide color gamut&amp;rdquo; such as 95%+ DCI-P3 coverage, but this can backfire for programming purposes.
In a Windows environment, if a wide color gamut monitor is used without applying the appropriate ICC profile (a color profile defined by the International Color Consortium), the syntax highlighting in VS Code, which is designated in standard sRGB (e.g., red or green warning colors), will be displayed in unnaturally vivid, oversaturated hues.
Because these intense colors strongly stimulate the eyes, it is highly recommended to either install the correct ICC profile from the OS display settings or switch the monitor&amp;rsquo;s OSD settings to &amp;ldquo;sRGB Emulation Mode.&amp;rdquo;&lt;/p>
&lt;p>The sequence diagram below illustrates the process of rendering eye-friendly colors once the correct ICC profile is applied.&lt;/p>
&lt;pre class="mermaid">
sequenceDiagram
participant OS as &amp;#34;Operating System&amp;#34;
participant LUT as &amp;#34;Color LUT (Look-Up Table)&amp;#34;
participant Mon as &amp;#34;Monitor Display&amp;#34;
participant Eye as &amp;#34;Programmer&amp;#39;s Eye&amp;#34;
OS-&amp;gt;&amp;gt;LUT: &amp;#34;Load Correct ICC Profile (e.g. sRGB)&amp;#34;
OS-&amp;gt;&amp;gt;LUT: &amp;#34;Apply Night Light Settings (3400K)&amp;#34;
LUT-&amp;gt;&amp;gt;Mon: &amp;#34;Adjust RGB Signal Output&amp;#34;
Mon-&amp;gt;&amp;gt;Eye: &amp;#34;Render Accurate, Desaturated Colors&amp;#34;
Eye--&amp;gt;&amp;gt;Eye: &amp;#34;Reduced Visual Cortical Strain&amp;#34;
&lt;/pre>
&lt;h2 id="32-software-countermeasures-flux--night-light">3.2 Software Countermeasures (f.lux / Night Light)
&lt;/h2>&lt;p>The easiest and most effective measure against blue light is software that dynamically changes the Color Temperature according to the time of day.&lt;/p>
&lt;ul>
&lt;li>Windows: &lt;strong>Night Light&lt;/strong>&lt;/li>
&lt;li>macOS: &lt;strong>Night Shift&lt;/strong>&lt;/li>
&lt;li>Third-party: &lt;strong>f.lux&lt;/strong>&lt;/li>
&lt;/ul>
&lt;p>Color temperature is expressed in Kelvin ($\text{K}$). Daytime sunlight is approximately $5500\text{K} \sim 6500\text{K}$ (bluish-white light), but if your eyes are constantly exposed to this, the secretion of &amp;ldquo;melatonin (sleep hormone)&amp;rdquo; in the pineal gland of the brain is suppressed.
After evening, by using these software tools to lower the color temperature down to $3400\text{K} \sim 1900\text{K}$ (warm orange to red), you can physically reduce the amount of blue light emitted. This not only keeps your circadian rhythm (body clock) normal but also prevents high-energy photons from reaching your eyeballs.&lt;/p>
&lt;hr>
&lt;h1 id="chapter-4-the-ultimate-hardware-solution-adopting-the-latest-gadgets">Chapter 4: The Ultimate Hardware Solution: Adopting the Latest Gadgets
&lt;/h1>&lt;p>If the measures explained so far fail to alleviate your fatigue, you need to invest in external gadgets to drastically change your environment.&lt;/p>
&lt;h2 id="41-bias-lighting-and-monitor-light-bars-screenbar">4.1 Bias Lighting and Monitor Light Bars (ScreenBar)
&lt;/h2>&lt;p>When you stare at a bright monitor in a dark room, a severe contrast occurs between the center of your field of vision (high brightness) and the periphery (low brightness). This is called &lt;strong>&amp;ldquo;Discomfort Glare.&amp;rdquo;&lt;/strong>
Under these conditions, the eyes fall into a contradictory state where they try to dilate the pupils to take in light while simultaneously trying to constrict them against the central glare, leading to severe fatigue of the iris muscles.&lt;/p>
&lt;p>The solution to this is &amp;ldquo;Bias Lighting.&amp;rdquo;
Particularly recommended are &amp;ldquo;monitor light bars&amp;rdquo; like the &lt;strong>BenQ ScreenBar&lt;/strong>.&lt;/p>
&lt;pre class="mermaid">
graph TD
A[&amp;#34;Dark Room Environment&amp;#34;] --&amp;gt; B[&amp;#34;High Brightness Contrast (Monitor vs Room)&amp;#34;]
B --&amp;gt; C[&amp;#34;Conflicting Pupil Constriction/Dilation&amp;#34;]
C --&amp;gt; D[&amp;#34;Severe Iris Muscle Fatigue&amp;#34;]
A --&amp;gt; E[&amp;#34;Install Monitor Light Bar (e.g., ScreenBar)&amp;#34;]
E --&amp;gt; F[&amp;#34;Asymmetrical Optical Design (No Glare on Screen)&amp;#34;]
F --&amp;gt; G[&amp;#34;Balanced Ambient Brightness&amp;#34;]
G --&amp;gt; H[&amp;#34;Relaxed Iris and Relieved Eye Strain&amp;#34;]
&lt;/pre>
&lt;p>The greatest feature of the ScreenBar is its &amp;ldquo;Asymmetrical Optical Design.&amp;rdquo; Through special reflectors and lenses, it does not shine light directly onto the monitor screen itself (preventing screen reflection and glare), and uniformly illuminates only the keyboard in front of you and the space behind the monitor. This dramatically mitigates the brightness difference (contrast ratio) across the entire field of vision, eliminating the burden on the eyes.&lt;/p>
&lt;h2 id="42-the-e-ink-display-paradigm-shift-dasung--boox">4.2 The E-Ink Display Paradigm Shift (Dasung &amp;amp; Boox)
&lt;/h2>&lt;p>For reading lengthy API references, technical books (PDFs), or code, the ultimate modern solution is using an &lt;strong>&amp;ldquo;E-Ink (electronic paper) display&amp;rdquo; as a secondary monitor&lt;/strong>.&lt;/p>
&lt;p>Unlike LCD or OLED, E-Ink does not have a self-emitting backlight. It displays text by applying voltage to move charged white and black pigment particles (such as titanium dioxide) inside capsules (electrophoresis), reflecting the ambient light.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Physical blue light emission: Zero&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Flicker associated with PWM or refresh rates: Absolutely zero&lt;/strong>&lt;/li>
&lt;/ul>
&lt;p>By placing an E-Ink monitor like the &lt;strong>Dasung Paperlike&lt;/strong> series (e.g., 25.3 inches) or &lt;strong>Onyx Boox Mira&lt;/strong> vertically as a dedicated text sub-monitor, you can read documents with the exact same sensation as reading printed paper.
While it has the drawback of drawing delay (low refresh rate), if limited strictly to the &amp;ldquo;reading of static text&amp;rdquo; in a programming environment, there is no device on earth kinder to the eyes.&lt;/p>
&lt;hr>
&lt;h1 id="chapter-5-ergonomics-and-operational-rules">Chapter 5: Ergonomics and Operational Rules
&lt;/h1>&lt;p>No matter how excellent the hardware you assemble is, it is meaningless if the human operating it has the wrong posture or rules.&lt;/p>
&lt;h2 id="51-fluid-dynamics-of-dry-eyes-and-line-of-sight-angle">5.1 Fluid Dynamics of Dry Eyes and Line of Sight Angle
&lt;/h2>&lt;p>Dry eyes are not just a discomfort of &amp;ldquo;eyes feeling dry.&amp;rdquo; When the tear film on the surface of the cornea is destroyed, light diffuses irregularly, blurring your vision, which leads to a vicious cycle of further eye strain (overworking the ciliary muscles).
The evaporation rate of tears is proportional to the surface area of the eyeball exposed to the air (palpebral fissure area).&lt;/p>
&lt;p>The ideal line of sight angle $\theta$ for monitor placement is considered to be $15^\circ \sim 20^\circ$ downward from the horizontal line.
When the horizontal distance from the center of the monitor to the eye is $d$, and the height difference between the monitor center and eye level is $h$, the following trigonometric function holds true:&lt;/p>
$$ \tan \theta = \frac{h}{d} $$&lt;p>For example, if the distance $d$ to the monitor is $60 \text{ cm}$ (a typical desk environment), to set $\theta = 15^\circ$:&lt;/p>
$$ h = 60 \times \tan(15^\circ) \approx 60 \times 0.267 = 16.02 \text{ cm} $$&lt;p>In other words, &lt;strong>ideally, the center of the monitor should be about $16 \text{ cm}$ below eye level.&lt;/strong>
By directing your line of sight slightly downward, your upper eyelids naturally drop, reducing the exposed area of the eyeball, which can dramatically prevent tear evaporation. Install a monitor arm (like Ergotron) and accurately set this height down to the millimeter.&lt;/p>
&lt;h2 id="52-strict-adherence-to-and-automation-of-the-global-standard-20-20-20-rule">5.2 Strict Adherence to and Automation of the Global Standard &amp;ldquo;20-20-20 Rule&amp;rdquo;
&lt;/h2>&lt;p>The &amp;ldquo;20-20-20 Rule&amp;rdquo; is a recovery method for digital device eye strain recommended by the American Academy of Ophthalmology (AAO) and ophthalmologists worldwide.&lt;/p>
&lt;p>&lt;strong>&amp;ldquo;Every 20 minutes, look at something 20 feet (about 6 meters) away for 20 seconds.&amp;rdquo;&lt;/strong>&lt;/p>
&lt;p>Through this simple action, the extremely contracted ciliary muscles are forcibly relaxed, the crystalline lens thins out, and the focus adjustment function is reset.
Because programmers often lose track of time when entering a flow state, the engineer-like solution is to build a mechanism that automatically enforces this rule.
Below is an example of an extremely simple script using Python&amp;rsquo;s &lt;code>tkinter&lt;/code> that forcibly displays a popup every 20 minutes.&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt"> 1
&lt;/span>&lt;span class="lnt"> 2
&lt;/span>&lt;span class="lnt"> 3
&lt;/span>&lt;span class="lnt"> 4
&lt;/span>&lt;span class="lnt"> 5
&lt;/span>&lt;span class="lnt"> 6
&lt;/span>&lt;span class="lnt"> 7
&lt;/span>&lt;span class="lnt"> 8
&lt;/span>&lt;span class="lnt"> 9
&lt;/span>&lt;span class="lnt">10
&lt;/span>&lt;span class="lnt">11
&lt;/span>&lt;span class="lnt">12
&lt;/span>&lt;span class="lnt">13
&lt;/span>&lt;span class="lnt">14
&lt;/span>&lt;span class="lnt">15
&lt;/span>&lt;span class="lnt">16
&lt;/span>&lt;span class="lnt">17
&lt;/span>&lt;span class="lnt">18
&lt;/span>&lt;span class="lnt">19
&lt;/span>&lt;span class="lnt">20
&lt;/span>&lt;span class="lnt">21
&lt;/span>&lt;span class="lnt">22
&lt;/span>&lt;span class="lnt">23
&lt;/span>&lt;span class="lnt">24
&lt;/span>&lt;span class="lnt">25
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-python" data-lang="python">&lt;span class="line">&lt;span class="cl">&lt;span class="kn">import&lt;/span> &lt;span class="nn">time&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="kn">import&lt;/span> &lt;span class="nn">tkinter&lt;/span> &lt;span class="k">as&lt;/span> &lt;span class="nn">tk&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="kn">from&lt;/span> &lt;span class="nn">tkinter&lt;/span> &lt;span class="kn">import&lt;/span> &lt;span class="n">messagebox&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">def&lt;/span> &lt;span class="nf">remind_20_20_20&lt;/span>&lt;span class="p">():&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="c1"># Hide the main window&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">root&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="n">tk&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">Tk&lt;/span>&lt;span class="p">()&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">root&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">withdraw&lt;/span>&lt;span class="p">()&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="k">while&lt;/span> &lt;span class="kc">True&lt;/span>&lt;span class="p">:&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="c1"># Wait for 20 minutes (1200 seconds)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">time&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">sleep&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="mi">20&lt;/span> &lt;span class="o">*&lt;/span> &lt;span class="mi">60&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="c1"># Display a warning dialog in the foreground&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">messagebox&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">showinfo&lt;/span>&lt;span class="p">(&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">title&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="s2">&amp;#34;20-20-20 Rule&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">message&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="s2">&amp;#34;Please look away from the screen and stare at something at least 6 meters away for 20 seconds!&lt;/span>&lt;span class="se">\n&lt;/span>&lt;span class="s2">(To relax your ciliary muscles)&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="c1"># 20 seconds for relaxation&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">time&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">sleep&lt;/span>&lt;span class="p">(&lt;/span>&lt;span class="mi">20&lt;/span>&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="k">if&lt;/span> &lt;span class="vm">__name__&lt;/span> &lt;span class="o">==&lt;/span> &lt;span class="s1">&amp;#39;__main__&amp;#39;&lt;/span>&lt;span class="p">:&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="c1"># Run in the background&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">remind_20_20_20&lt;/span>&lt;span class="p">()&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>By registering such a script at startup or running it via the OS standard task scheduler/Cron, you can integrate a mandatory recovery cycle into your daily life.&lt;/p>
&lt;hr>
&lt;h1 id="conclusion-eye-strain-countermeasures-as-an-investment-in-the-future">Conclusion: Eye Strain Countermeasures as an Investment in the Future
&lt;/h1>&lt;p>Our careers as software engineers will last for decades. What supports that career is not an expensive keyboard or the latest CPU, but undeniably our own &amp;ldquo;eyes&amp;rdquo; and &amp;ldquo;brain.&amp;rdquo;&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Understand the physical load of light energy ($E = hc/\lambda$) and focus adjustment.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Introduce a flicker-free (DC dimming) and high refresh rate monitor.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Optimize the relative contrast of the environment with bias lighting such as a ScreenBar.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Consider an E-Ink monitor as the ultimate text viewing device.&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Create an optimal line of sight angle based on $\tan \theta = h/d$ using a monitor arm, and systemize the &amp;ldquo;20-20-20 Rule.&amp;rdquo;&lt;/strong>&lt;/li>
&lt;/ol>
&lt;p>While these measures may involve temporary expenses and effort, they are arguably the most cost-effective &amp;ldquo;technical investments&amp;rdquo; to extend the healthy lifespan of your eyes and maximize your lifelong productivity and QOL (Quality of Life). Reevaluate your development environment right now and implement some compassion for your eyes.&lt;/p></description></item><item><title>Techniques for Resolving GPU Memory Shortages in AI Development (CPU Offloading, etc.)</title><link>http://kenji.blog/en/p/ai-gpu-vram-optimization-cpu-offloading/</link><pubDate>Fri, 11 Sep 2026 01:00:00 +0900</pubDate><guid>http://kenji.blog/en/p/ai-gpu-vram-optimization-cpu-offloading/</guid><description>&lt;img src="http://kenji.blog/p/ai-gpu-vram-optimization-cpu-offloading/img/eyecatch.jpg" alt="Featured image of post Techniques for Resolving GPU Memory Shortages in AI Development (CPU Offloading, etc.)" />&lt;h1 id="introduction-ai-development-and-the-wall-of-vram">Introduction: AI Development and &amp;ldquo;The Wall of VRAM&amp;rdquo;
&lt;/h1>&lt;p>In recent years, generative AI technologies such as Large Language Models (LLMs) and Diffusion Models have achieved rapid development. However, when trying to train (fine-tune) or perform inference with these cutting-edge AI models in a local environment, a very physical barrier that many developers and researchers face is the &lt;strong>&amp;ldquo;shortage of GPU memory (VRAM)&amp;rdquo;&lt;/strong>.&lt;/p>
&lt;p>Even a high-end consumer GPU like the NVIDIA GeForce RTX 4090 has a maximum VRAM of 24GB, making it completely impossible to load a massive model like Llama 3 70B as it is. Data center GPUs like the H100 (80GB) or B200 (192GB) are extremely expensive and not easily accessible to individuals or small teams. If we cannot break through this &amp;ldquo;Wall of VRAM,&amp;rdquo; we cannot even touch the most advanced models.&lt;/p>
&lt;p>In this article, we will thoroughly explain advanced techniques from both inference and training perspectives to break down this physical constraint of limited VRAM through software and hardware architectural ingenuity. Let&amp;rsquo;s delve deep into CPU offloading, KV cache optimization, Gradient Checkpointing, and the latest Unified Memory architectures, interspersed with mathematical formulas and diagrams. By reading this article, you will gain a deep understanding of VRAM behavior and practical knowledge for handling massive models with limited resources.&lt;/p>
&lt;hr>
&lt;h1 id="1-anatomy-of-vram-consumption-by-ai-models-inference--training">1. Anatomy of VRAM Consumption by AI Models (Inference &amp;amp; Training)
&lt;/h1>&lt;p>The first step to resolving VRAM shortages is to accurately grasp &amp;ldquo;what&amp;rdquo; is consuming &amp;ldquo;how much&amp;rdquo; memory from a micro perspective. By treating it not as a black box, but by accurately estimating it using mathematical formulas, we can select appropriate optimization techniques.&lt;/p>
&lt;h2 id="11-memory-calculation-for-model-parameters-weights">1.1 Memory Calculation for Model Parameters (Weights)
&lt;/h2>&lt;p>The basic amount of memory consumed by the parameters (Weights) that make up an AI model is determined by the total number of parameters in the model and the data type (Precision) used to represent them.&lt;/p>
&lt;p>The data types commonly used in deep learning and their byte size per parameter ($B$) are as follows:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>FP32 (Single Precision Floating-Point):&lt;/strong> 4 bytes (Standard precision for training)&lt;/li>
&lt;li>&lt;strong>FP16 / BF16 (Half Precision Floating-Point):&lt;/strong> 2 bytes (Common for inference and mixed precision training)&lt;/li>
&lt;li>&lt;strong>INT8 (8-bit Integer):&lt;/strong> 1 byte (Quantized models)&lt;/li>
&lt;li>&lt;strong>INT4 (4-bit Integer Quantization):&lt;/strong> 0.5 bytes (Extreme quantization like GPTQ, AWQ, GGUF)&lt;/li>
&lt;/ul>
&lt;p>Let $P$ be the total number of parameters in the model. The base memory amount $M_{weights}$ occupied by the weights themselves is expressed by the following formula:&lt;/p>
$$ M_{weights} = P \times B $$&lt;p>For example, when loading Meta&amp;rsquo;s &amp;ldquo;Llama 3 8B&amp;rdquo; model (approx. 8 billion parameters) in FP16 (half precision), the calculation is as follows:&lt;/p>
$$ M_{weights} = 8,000,000,000 \times 2 \text{ bytes} \approx 16,000,000,000 \text{ bytes} \approx 16 \text{ GB} $$&lt;p>In other words, purely loading the model&amp;rsquo;s weights into the GPU consumes 16GB of VRAM. With an RTX 3060 (12GB), an Out of Memory (OOM) error will occur at this point. However, if the model is quantized to INT4, it becomes $8 \times 0.5 = 4 \text{ GB}$, making it easily loadable.&lt;/p>
&lt;h2 id="12-memory-consumption-during-inference-the-growth-of-kv-cache">1.2 Memory Consumption During Inference: The Growth of KV Cache
&lt;/h2>&lt;p>In LLM inference (especially autoregressive text generation), what heavily pressures VRAM as much as or more than the weights is the &lt;strong>KV Cache (Key-Value Cache)&lt;/strong>.
In the Transformer architecture, to avoid recalculating information for tokens that were already generated and processed, the Key and Value tensors in each attention layer are continuously cached in VRAM. This improves computational speed (Compute), but as the context length (input prompt length + generated length) grows, memory consumption increases linearly and explosively.&lt;/p>
&lt;p>The amount of KV cache memory consumed when processing a single token, $M_{kv\_token}$, is strictly calculated based on the model&amp;rsquo;s architecture by the following formula:&lt;/p>
$$ M_{kv\_token} = 2 \times N_{layers} \times N_{heads\_kv} \times D_{head} \times B $$&lt;p>Here, each variable has the following meaning:&lt;/p>
&lt;ul>
&lt;li>$2$ : Because there are two tensors, Key and Value&lt;/li>
&lt;li>$N_{layers}$ : Number of Transformer layers&lt;/li>
&lt;li>$N_{heads\_kv}$ : Number of KV attention heads (In GQA: Grouped Query Attention, this is fewer than the standard number of heads)&lt;/li>
&lt;li>$D_{head}$ : Dimensionality of each head (Usually, hidden layer dimension $D_{model} / N_{heads}$)&lt;/li>
&lt;li>$B$ : Number of bytes for the data type (2 for FP16)&lt;/li>
&lt;/ul>
&lt;p>The total KV cache amount $M_{kv\_total}$ is this multiplied by the sequence length ($L_{seq}$) and the batch size ($BatchSize$).&lt;/p>
$$ M_{kv\_total} = M_{kv\_token} \times L_{seq} \times BatchSize $$&lt;p>&lt;strong>Example: For Llama 2 7B&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>$N_{layers} = 32$&lt;/li>
&lt;li>$N_{heads\_kv} = 32$ (For MHA)&lt;/li>
&lt;li>$D_{head} = 128$&lt;/li>
&lt;li>FP16 ($B=2$)&lt;/li>
&lt;li>Batch Size 1, Sequence Length 8192 (8K Context)&lt;/li>
&lt;/ul>
$$ M_{kv\_total} = 2 \times 32 \times 32 \times 128 \times 2 \times 8192 \times 1 = 4,294,967,296 \text{ bytes} \approx 4 \text{ GB} $$&lt;p>If we extend the context to 32K (32768 tokens), the KV cache alone will consume about 16GB. If we increase the batch size to 4, it will be 64GB. The fact that it demands significantly more VRAM than the size of the model itself is a major challenge during inference.&lt;/p>
&lt;h2 id="13-memory-consumption-during-training-optimizers-gradients-and-activations">1.3 Memory Consumption During Training: Optimizers, Gradients, and Activations
&lt;/h2>&lt;p>Compared to inference, training a model (pre-training or fine-tuning) consumes vastly more VRAM. This is because, in addition to a simple forward pass, it is necessary to hold information for backpropagation. Training memory is primarily composed of the following four elements:&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Model Weights:&lt;/strong> Similar to inference, but mixed precision training may require holding both FP16 and FP32 (master weights).&lt;/li>
&lt;li>&lt;strong>Gradients:&lt;/strong> The gradients for each parameter calculated during backpropagation. 2 bytes per parameter for FP16.&lt;/li>
&lt;li>&lt;strong>Optimizer States:&lt;/strong> Advanced optimizers like AdamW hold the first moment (Momentum) and second moment (Variance) for each parameter. To maintain training stability, these are typically held in FP32 (4 bytes). Thus, the two moments consume $4 + 4 = 8$ bytes per parameter.&lt;/li>
&lt;li>&lt;strong>Activations:&lt;/strong> To calculate gradients during backpropagation, the output of each layer from the forward pass (intermediate states) must be kept in memory. This heavily depends on the batch size and sequence length, and becomes exceptionally large.&lt;/li>
&lt;/ol>
&lt;p>In summary, in Mixed Precision Training using a standard Adam optimizer, &lt;strong>about 16 to 20 bytes&lt;/strong> (Master weights 4 + FP16 weights 2 + Gradients 2 + Optimizer 8 + α) of memory is required per parameter.&lt;/p>
$$ M_{train\_param} \approx P \times 16 \text{ bytes} $$&lt;p>To train a 7B (7 billion parameter) model, parameter-related elements alone will require $7B \times 16 = 112 \text{ GB}$, and with activations added, it equates to needing over 140GB of VRAM. Executing this on 24GB of VRAM requires the aggressive optimization techniques explained in the subsequent chapters.&lt;/p>
&lt;hr>
&lt;h1 id="2-vram-saving-techniques-during-inference">2. VRAM Saving Techniques During Inference
&lt;/h1>&lt;p>To run massive models during inference, many software technologies have been developed that transcend hardware boundaries.&lt;/p>
&lt;h2 id="21-cpu-offloading-and-layer-splitting">2.1 CPU Offloading and Layer Splitting
&lt;/h2>&lt;p>When a massive model cannot entirely fit on a single or multiple GPUs, the technique of placing parts of the model in system memory (CPU RAM) and transferring them to the GPU only when necessary as computation progresses is known as &lt;strong>CPU Offloading&lt;/strong>. &lt;code>llama.cpp&lt;/code> and Hugging Face&amp;rsquo;s &lt;code>Accelerate&lt;/code> support this feature.&lt;/p>
&lt;pre class="mermaid">
graph TD
A[&amp;#34;System RAM (DDR4 / DDR5)&amp;#34;] --&amp;gt; B[&amp;#34;GPU VRAM (GDDR6X)&amp;#34;]
B[&amp;#34;GPU VRAM (GDDR6X)&amp;#34;] --&amp;gt; C[&amp;#34;Tensor Cores (Compute)&amp;#34;]
subgraph &amp;#34;Layer Splitting and Offloading&amp;#34;
D[&amp;#34;Lower Layers 1-15 (GPU Pinned)&amp;#34;]
E[&amp;#34;Upper Layers 16-32 (CPU Offloaded)&amp;#34;]
end
E[&amp;#34;Upper Layers 16-32 (CPU Offloaded)&amp;#34;] -.-&amp;gt; B[&amp;#34;GPU VRAM (GDDR6X)&amp;#34;]
&lt;/pre>
&lt;p>&lt;strong>Mechanism and Challenges:&lt;/strong>
Since Transformer models have a structure where layers are stacked in series, the computation of one layer must finish before the next layer&amp;rsquo;s computation can begin. Leveraging this, only the layers that fit in the GPU (e.g., Layers 1-15) are pinned residently in the VRAM, while the remaining layers (Layers 16-32) are placed in the large-capacity but slower CPU RAM. During inference, when the computation for up to layer 15 is complete, the weights for layer 16 are transferred (copied) from the CPU to the GPU via the PCIe bus, and computation is executed on the GPU.&lt;/p>
&lt;p>However, &lt;strong>PCIe bandwidth becomes a severe bottleneck&lt;/strong>. The theoretical maximum bandwidth of PCIe 4.0 x16 is 32GB/s (unidirectional), but compared to the internal bandwidth of modern GPU VRAM (e.g., RTX 4090&amp;rsquo;s GDDR6X is 1008GB/s, H100&amp;rsquo;s HBM3 is over 3TB/s), it is two orders of magnitude slower. Therefore, heavily relying on CPU offloading dramatically decreases inference speed (Tokens per Second).
To minimize speed degradation, the practical key is to place as many layers as possible on the GPU (maximizing GPU Layers) and minimize the number of offloaded layers.&lt;/p>
&lt;h2 id="22-kv-cache-quantization-and-pagedattention">2.2 KV Cache Quantization and PagedAttention
&lt;/h2>&lt;p>For the KV cache, which is the primary culprit of VRAM consumption during inference, two powerful optimizations have been implemented.&lt;/p>
&lt;p>&lt;strong>1. KV Cache Quantization:&lt;/strong>
This is a technique where not only the model weights, but the dynamically generated KV cache itself is quantized to INT8, INT4, or even FP8 before being saved in VRAM. This can reduce the size of the KV cache by half to a quarter. Modern inference engines (like vLLM and llama.cpp) have incorporated this feature, achieving significant VRAM savings while minimizing precision loss.&lt;/p>
&lt;p>&lt;strong>2. PagedAttention:&lt;/strong>
Applying the concept of OS virtual memory &amp;ldquo;paging&amp;rdquo; to the KV cache resulted in &lt;strong>PagedAttention&lt;/strong>, introduced in the vLLM inference engine. In traditional inference engines, contiguous VRAM regions were pre-allocated based on the set maximum sequence length. As a result, when actual inputs were shorter, fragmentation and wasted memory occurred, sometimes wasting over 60% of VRAM.&lt;/p>
&lt;p>PagedAttention divides the KV cache into fixed-size blocks (pages), making it possible to store them distributed across non-contiguous physical memory spaces. This brings memory waste to near zero (limited only to internal fragmentation) and allows the batch size to be significantly increased with the same VRAM capacity.&lt;/p>
&lt;pre class="mermaid">
graph LR
A[&amp;#34;Logical KV Cache&amp;#34;] --&amp;gt; B[&amp;#34;Physical VRAM Blocks&amp;#34;]
A1[&amp;#34;Token 1, 2, 3, 4&amp;#34;] --&amp;gt; B3[&amp;#34;Block 3 (Allocated)&amp;#34;]
A2[&amp;#34;Token 5, 6, 7, 8&amp;#34;] --&amp;gt; B1[&amp;#34;Block 1 (Allocated)&amp;#34;]
A3[&amp;#34;Future Tokens...&amp;#34;] -.-&amp;gt; B2[&amp;#34;Block 2 (Free)&amp;#34;]
&lt;/pre>
&lt;h2 id="23-flashattention-breaking-the-memory-complexity-of-attention-computation">2.3 FlashAttention: Breaking the Memory Complexity of Attention Computation
&lt;/h2>&lt;p>VRAM shortage is caused not only by the memory needed to store data, but also by the shortage of &amp;ldquo;temporary workspace&amp;rdquo; during computation. The standard Transformer Self-Attention mechanism needs to materialize a massive $N \times N$ attention matrix on VRAM for a sequence length $N$. This has a memory complexity of $O(N^2)$ and is a major cause of OOM with long contexts.&lt;/p>
&lt;p>What solved this is &lt;strong>FlashAttention&lt;/strong> (and FlashAttention-2, 3).
FlashAttention is an algorithm that is aware of the GPU hardware architecture (the hierarchical structure of large but slow HBM and extremely small but ultra-fast SRAM). Using a technique called Tiling, it loads data into SRAM block by block and completes the attention computation there, thereby completely avoiding the process of writing the $N \times N$ matrix to HBM (VRAM).&lt;/p>
&lt;p>As a result, the memory complexity of attention layers drops dramatically from $O(N^2)$ to $O(N)$ (proportional to sequence length), significantly easing constraints on context length.&lt;/p>
&lt;h2 id="24-the-rise-of-unified-memory-and-apple-silicon">2.4 The Rise of Unified Memory and Apple Silicon
&lt;/h2>&lt;p>Approaching this problem from the very foundation of PC architecture is the &lt;strong>Unified Memory Architecture (UMA)&lt;/strong> adopted by Apple Silicon (M1/M2/M3/M4 series Max and Ultra) and some modern APUs (like AMD Strix Point).&lt;/p>
&lt;p>In these architectures, the CPU and GPU on the motherboard share the exact same physical memory (e.g., up to 192GB of LPDDR5). Therefore, the concept of &amp;ldquo;slow data transfer from CPU to GPU via PCIe&amp;rdquo; physically does not exist.&lt;/p>
&lt;pre class="mermaid">
graph TD
subgraph &amp;#34;Unified Memory Architecture (e.g. Apple Silicon)&amp;#34;
A[&amp;#34;CPU Cores&amp;#34;] &amp;lt;--&amp;gt; C[&amp;#34;Shared Memory Controller&amp;#34;]
B[&amp;#34;GPU Cores / Neural Engine&amp;#34;] &amp;lt;--&amp;gt; C[&amp;#34;Shared Memory Controller&amp;#34;]
C[&amp;#34;Shared Memory Controller&amp;#34;] &amp;lt;--&amp;gt; D[&amp;#34;Unified Memory Pool (e.g. 192GB)&amp;#34;]
end
&lt;/pre>
&lt;p>The greatest advantage of this architecture is that there is no distinct wall of VRAM; almost the entire system memory can be directly used to load massive LLMs. With a Mac Studio possessing 192GB of unified memory, it is possible to load massive models of the 70B class or larger (such as Grok-1) onto a single device without quantization and infer at high speed. The memory access bandwidth also reaches 800GB/s on the M2 Ultra, boasting speeds comparable to consumer discrete GPUs. It is a very powerful approach that solves the dilemma of &amp;ldquo;memory capacity&amp;rdquo; and &amp;ldquo;bandwidth&amp;rdquo; at the hardware level.&lt;/p>
&lt;hr>
&lt;h1 id="3-vram-saving-techniques-during-training-fine-tuning">3. VRAM Saving Techniques During Training (Fine-Tuning)
&lt;/h1>&lt;p>During training, which demands even more VRAM than inference, many breakthroughs have also been made. To perform fine-tuning with limited resources, combining the following technologies is essential.&lt;/p>
&lt;h2 id="31-gradient-checkpointing">3.1 Gradient Checkpointing
&lt;/h2>&lt;p>In deep learning backpropagation, to calculate gradients, the intermediate outputs (Activations) of all layers from the forward pass must be kept in memory. When sequence lengths and batch sizes increase, this activation memory begins to dominate VRAM.&lt;/p>
&lt;p>&lt;strong>Gradient Checkpointing (or Activation Recomputation)&lt;/strong> is an ingenious technique that trades memory capacity for computation time (Compute).
Instead of saving all intermediate outputs in memory, it saves only the outputs of specific layers (checkpoints). During backpropagation, if an unsaved intermediate value is needed, &lt;strong>it restores the value by recalculating the forward pass from the nearest saved checkpoint&lt;/strong>.&lt;/p>
&lt;p>Computational overhead increases by about 20-30%, lengthening total training time, but it drastically reduces VRAM consumption by activations from $O(N)$ ($N$ being the number of layers) down to $O(\sqrt{N})$. In current large-scale model training, it is an essential setting to the point where one can say you cannot even start without it.&lt;/p>
&lt;h2 id="32-lora-and-qlora-low-rank-adaptation">3.2 LoRA and QLoRA (Low-Rank Adaptation)
&lt;/h2>&lt;p>The star player that fundamentally solved VRAM shortages is &lt;strong>LoRA&lt;/strong>, a representative of PEFT (Parameter-Efficient Fine-Tuning).&lt;/p>
&lt;p>The original massive weight matrix of the model, $W_0 \in \mathbb{R}^{d \times k}$, is frozen and not trained. Instead, two very small, low-rank matrices $A \in \mathbb{R}^{r \times k}$ and $B \in \mathbb{R}^{d \times r}$ are introduced in parallel, and only this $A$ and $B$ are trained. (Here, the rank $r$ is a small value such that $r \ll d, k$).&lt;/p>
$$ W_{adapted} = W_0 + \Delta W = W_0 + B A $$&lt;p>By doing this, the number of parameters to be trained drops to less than 1% (sometimes less than 0.1%) of the original, and correspondingly, the &amp;ldquo;gradients&amp;rdquo; and &amp;ldquo;optimizer states&amp;rdquo; that consumed a massive amount of memory decrease sharply to less than 1%.&lt;/p>
&lt;p>Taking this to its absolute limit is &lt;strong>QLoRA (Quantized LoRA)&lt;/strong>.
In QLoRA, the base model weights $W_0$ are heavily quantized to 4-bit (NF4: NormalFloat4 format) and loaded into VRAM. Then, LoRA&amp;rsquo;s small matrices $A, B$ are trained in BF16 (16-bit) to maintain computational precision.
While reducing the base model&amp;rsquo;s VRAM size to a quarter of its original size through 4-bit quantization, it also uses a technique called &lt;strong>Paged Optimizers&lt;/strong> to automatically evict (offload) the optimizer states to CPU RAM temporarily when VRAM is about to be exhausted. Thanks to this, fine-tuning of ultra-massive models like Llama 3 70B became possible even on a single GPU with 24GB VRAM (like an RTX 4090).&lt;/p>
&lt;h2 id="33-deepspeed-zero-and-offloading">3.3 DeepSpeed ZeRO and Offloading
&lt;/h2>&lt;p>In an environment using multiple GPUs (Multi-GPU), simple Data Parallelism does not resolve the VRAM issue. Because each GPU holds a copy of the entire model, the limits of individual VRAM capacities cannot be exceeded.&lt;/p>
&lt;p>&lt;strong>ZeRO (Zero Redundancy Optimizer)&lt;/strong>, developed by Microsoft&amp;rsquo;s &lt;strong>DeepSpeed&lt;/strong> library, is a technique that thoroughly partitions (shards) model parameters, gradients, and optimizer states across multiple GPUs. This allows the &amp;ldquo;total value&amp;rdquo; of the VRAM across multiple GPUs to be handled like one massive memory pool.&lt;/p>
&lt;pre class="mermaid">
graph TD
subgraph &amp;#34;ZeRO Stage 3 (Parameter Partitioning)&amp;#34;
A[&amp;#34;GPU 0&amp;#34;] --&amp;gt; D[&amp;#34;Partition 0 (Stores 1/3 of Weights/Grads/Opts)&amp;#34;]
B[&amp;#34;GPU 1&amp;#34;] --&amp;gt; E[&amp;#34;Partition 1 (Stores 1/3 of Weights/Grads/Opts)&amp;#34;]
C[&amp;#34;GPU 2&amp;#34;] --&amp;gt; F[&amp;#34;Partition 2 (Stores 1/3 of Weights/Grads/Opts)&amp;#34;]
end
D[&amp;#34;Partition 0 (Stores 1/3 of Weights/Grads/Opts)&amp;#34;] &amp;lt;--&amp;gt; E[&amp;#34;Partition 1 (Stores 1/3 of Weights/Grads/Opts)&amp;#34;]
E[&amp;#34;Partition 1 (Stores 1/3 of Weights/Grads/Opts)&amp;#34;] &amp;lt;--&amp;gt; F[&amp;#34;Partition 2 (Stores 1/3 of Weights/Grads/Opts)&amp;#34;]
&lt;/pre>
&lt;ul>
&lt;li>&lt;strong>ZeRO Stage 1:&lt;/strong> Partitions optimizer states to each GPU.&lt;/li>
&lt;li>&lt;strong>ZeRO Stage 2:&lt;/strong> Partitions gradients to each GPU as well.&lt;/li>
&lt;li>&lt;strong>ZeRO Stage 3:&lt;/strong> Partitions the model parameters (weights) themselves to each GPU.&lt;/li>
&lt;/ul>
&lt;p>Furthermore, using a feature called &lt;strong>ZeRO-Offload&lt;/strong>, the update calculations of optimizer states and gradients partitioned by ZeRO can be &lt;strong>offloaded to CPU memory&lt;/strong> and executed on the host CPU instead of the GPU. This minimizes the burden on GPU VRAM to the absolute limit, enabling the training of massive models even in constrained GPU environments. Since the calculations are done on the CPU and results are returned to the GPU via PCIe, training speed decreases, but you can avoid the worst-case scenario where &amp;ldquo;training crashes due to out of memory.&amp;rdquo;&lt;/p>
&lt;hr>
&lt;h1 id="4-implementation-example-hugging-face-accelerate-and-deepspeed">4. Implementation Example: Hugging Face Accelerate and DeepSpeed
&lt;/h1>&lt;p>Finally, here are simple examples showing how CPU offloading and VRAM optimization are actually implemented using Python code.&lt;/p>
&lt;h2 id="41-automatic-offloading-with-hugging-face-device_mapauto">4.1 Automatic Offloading with Hugging Face &lt;code>device_map=&amp;quot;auto&amp;quot;&lt;/code>
&lt;/h2>&lt;p>By using Hugging Face&amp;rsquo;s &lt;code>transformers&lt;/code> and &lt;code>accelerate&lt;/code> libraries, layers are automatically partitioned between the GPU and CPU when loading the model.&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt"> 1
&lt;/span>&lt;span class="lnt"> 2
&lt;/span>&lt;span class="lnt"> 3
&lt;/span>&lt;span class="lnt"> 4
&lt;/span>&lt;span class="lnt"> 5
&lt;/span>&lt;span class="lnt"> 6
&lt;/span>&lt;span class="lnt"> 7
&lt;/span>&lt;span class="lnt"> 8
&lt;/span>&lt;span class="lnt"> 9
&lt;/span>&lt;span class="lnt">10
&lt;/span>&lt;span class="lnt">11
&lt;/span>&lt;span class="lnt">12
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-python" data-lang="python">&lt;span class="line">&lt;span class="cl">&lt;span class="kn">from&lt;/span> &lt;span class="nn">transformers&lt;/span> &lt;span class="kn">import&lt;/span> &lt;span class="n">AutoModelForCausalLM&lt;/span>&lt;span class="p">,&lt;/span> &lt;span class="n">AutoTokenizer&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="n">model_id&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="s2">&amp;#34;meta-llama/Llama-2-13b-hf&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="c1"># With device_map=&amp;#34;auto&amp;#34;, parts that don&amp;#39;t fit into VRAM are offloaded to CPU RAM&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="c1"># load_in_8bit=True quantizes weights to 8-bit, saving even more memory&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="n">model&lt;/span> &lt;span class="o">=&lt;/span> &lt;span class="n">AutoModelForCausalLM&lt;/span>&lt;span class="o">.&lt;/span>&lt;span class="n">from_pretrained&lt;/span>&lt;span class="p">(&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">model_id&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">device_map&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="s2">&amp;#34;auto&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">load_in_8bit&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="kc">True&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="n">offload_folder&lt;/span>&lt;span class="o">=&lt;/span>&lt;span class="s2">&amp;#34;offload_dir&amp;#34;&lt;/span> &lt;span class="c1"># If even CPU RAM isn&amp;#39;t enough, it can offload to disk (SSD)&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">)&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>When you execute this code, the &lt;code>accelerate&lt;/code> library behind the scenes analyzes the free capacity of the system&amp;rsquo;s VRAM and CPU RAM, and dispatches the layers in the most optimal configuration.&lt;/p>
&lt;h2 id="42-deepspeed-cpu-offloading-configuration-zero-2">4.2 DeepSpeed CPU Offloading Configuration (ZeRO-2)
&lt;/h2>&lt;p>This is an example of a configuration file (JSON) to enable CPU offloading with DeepSpeed during training.&lt;/p>
&lt;div class="highlight">&lt;div class="chroma">
&lt;table class="lntable">&lt;tr>&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code>&lt;span class="lnt"> 1
&lt;/span>&lt;span class="lnt"> 2
&lt;/span>&lt;span class="lnt"> 3
&lt;/span>&lt;span class="lnt"> 4
&lt;/span>&lt;span class="lnt"> 5
&lt;/span>&lt;span class="lnt"> 6
&lt;/span>&lt;span class="lnt"> 7
&lt;/span>&lt;span class="lnt"> 8
&lt;/span>&lt;span class="lnt"> 9
&lt;/span>&lt;span class="lnt">10
&lt;/span>&lt;span class="lnt">11
&lt;/span>&lt;span class="lnt">12
&lt;/span>&lt;span class="lnt">13
&lt;/span>&lt;span class="lnt">14
&lt;/span>&lt;span class="lnt">15
&lt;/span>&lt;span class="lnt">16
&lt;/span>&lt;span class="lnt">17
&lt;/span>&lt;span class="lnt">18
&lt;/span>&lt;span class="lnt">19
&lt;/span>&lt;span class="lnt">20
&lt;/span>&lt;/code>&lt;/pre>&lt;/td>
&lt;td class="lntd">
&lt;pre tabindex="0" class="chroma">&lt;code class="language-json" data-lang="json">&lt;span class="line">&lt;span class="cl">&lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;fp16&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;enabled&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="kc">true&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">},&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;zero_optimization&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;stage&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="mi">2&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;offload_optimizer&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="p">{&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;device&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="s2">&amp;#34;cpu&amp;#34;&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;pin_memory&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="kc">true&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">},&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;allgather_partitions&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="kc">true&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;allgather_bucket_size&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="mf">2e8&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;overlap_comm&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="kc">true&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;reduce_scatter&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="kc">true&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;reduce_bucket_size&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="mf">2e8&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;contiguous_gradients&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="kc">true&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="p">},&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;train_batch_size&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="mi">16&lt;/span>&lt;span class="p">,&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl"> &lt;span class="nt">&amp;#34;gradient_accumulation_steps&amp;#34;&lt;/span>&lt;span class="p">:&lt;/span> &lt;span class="mi">4&lt;/span>
&lt;/span>&lt;/span>&lt;span class="line">&lt;span class="cl">&lt;span class="p">}&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/td>&lt;/tr>&lt;/table>
&lt;/div>
&lt;/div>&lt;p>In this configuration, by setting &lt;code>offload_optimizer&lt;/code> to &lt;code>&amp;quot;cpu&amp;quot;&lt;/code>, the state retention and update calculations for optimizers (like Adam) that consume vast amounts of VRAM are executed on the system&amp;rsquo;s CPU. This allows the GPU VRAM to focus exclusively on its most critical task: the forward and backward pass calculations of the model. By setting &lt;code>pin_memory: true&lt;/code>, page faults are prevented, and PCIe transfers between CPU and GPU are accelerated as much as possible.&lt;/p>
&lt;hr>
&lt;h1 id="conclusion">Conclusion
&lt;/h1>&lt;p>GPU memory shortage (Out of Memory) in AI development is an eternal challenge that will continue to follow developers as models scale up. However, by properly combining a deep understanding of hardware (architecture) with software and algorithm optimization techniques—as explained in this article—it becomes possible to perform inference and training for massive models in local environments, which might at first seem impossible.&lt;/p>
&lt;p>&lt;strong>Summary of Countermeasures During Inference:&lt;/strong>&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Quantization (INT4 / INT8 / FP8):&lt;/strong> Drastically compress the model&amp;rsquo;s size itself to reduce VRAM footprint.&lt;/li>
&lt;li>&lt;strong>CPU Offloading:&lt;/strong> Escape layers that do not fit in VRAM to system memory (a trade-off with speed drops due to PCIe bandwidth).&lt;/li>
&lt;li>&lt;strong>KV Cache Optimization:&lt;/strong> Secure context length using paging (PagedAttention), cache quantization, and FlashAttention.&lt;/li>
&lt;li>&lt;strong>Utilizing Unified Memory:&lt;/strong> Utilize UMAs like Apple Silicon to use large-capacity memory directly for inference.&lt;/li>
&lt;/ol>
&lt;p>&lt;strong>Summary of Countermeasures During Training:&lt;/strong>&lt;/p>
&lt;ol>
&lt;li>&lt;strong>PEFT (LoRA / QLoRA):&lt;/strong> Limit the parameters to be trained and aggressively quantize the base model.&lt;/li>
&lt;li>&lt;strong>Gradient Checkpointing:&lt;/strong> Discard intermediate forward pass outputs and recalculate them during the backward pass to suppress VRAM consumption in exchange for computation time.&lt;/li>
&lt;li>&lt;strong>ZeRO &amp;amp; CPU Offloading (DeepSpeed):&lt;/strong> Partition optimizer states and gradients across multiple GPUs, or offload them to CPU memory to break through VRAM limits.&lt;/li>
&lt;/ol>
&lt;p>By fully leveraging these advanced technologies, let&amp;rsquo;s extract the maximum AI development performance within limited hardware resources. In this rapidly advancing field, we can expect the emergence of new memory-saving algorithms in the future. Regularly checking the latest trends in libraries and incorporating them into your implementations will be key.&lt;/p></description></item></channel></rss>