• Nobody should use alloca. If you must allocate a buffer on the stack, use a VLA, which is standardized, has proper scope-bound lifetimes, and a type that remembers the exact size. Yes, I know MSVC does not upport it. Don't use this compiler. (where credit is due: MSVC had stack probing a lot ealier than GCC and clang, and clang was very late).

    With gcc, you get stack probing with -fstack-clash-protection, which is similar to _chkstk but GCC inlines the stack probes.

    VLA got a bad name because of stack clash attacks, but without stack clash protection these attacks can appear also without VLAs (and the first such attacks actually exploited fixed-size arrays), and if you activate this protection there is IMHO not much reason to avoid VLAs.

    If you need a small variably-sized buffers, VLAs are almost always superior to any alternative. alloca is worse in every way (see above), a regular array with worst-case bound increases stack use relative to a VLA and does not encode the correct dynamic size which makes bounds checking weaker, and moving the buffer to the heap is slower and complicates the code.

    If you can not properly account for the sizes of the things you put on your stack and worry about VLAs exceeding the limit (but again, regular arrays with worst-case size increase stack usage compared to VLAs), on GCC you can use -Wvla-larger-than to make sure the size of each VLA stays bounded.

  • One of those functions that isn't really implementable in standard C, requiring either compiler support, or being written in straight Assembly for stack registers manipulation, one of those "micro runtime" features for C.

    From UNIX 7th edition all the way up to C99, when VLAs where introduced, only to be made optional in C11, and the C23 update still doesn't support automatic VLAs, only for function parameters, thus the point stands.

  • So what if several functions that use less than 4KB each call each other before using the stack variables in a way that the first access skips over one page?
    • A function call causes the return address to be pushed onto the stack, thus accessing the stack below the adjusted stack pointer address.

      On compiler generated x86 code, the base pointer register will quickly follow when entering the target function.

  • One thing that scares me a little is whether there are younger developers, say, 25-40, who can and want to pick up the mantle of Windows internals gurus.

    I mean, Chen has decades of winternals in his head. Microsoft has been gutting their staff for years now. When the Petzold/Chen generation hang up their spurs, does Microsoft still have a critical mass of people who understand Windows from the metal up?

    • Why do you think Windows is currently such a mess?

      Microsoft new blood has been educated on Macs and ChromeOS, even if they do games it is most likely consoles.

      On WinUI community calls you usually would get puzzled faces when the Q&A touched when would WinUI be able to do "insert basic Win32/Forms/WPF" feature.

      Management apparently doesn't care they actually understand Windows, or get the required trainings to meet the quality of their predecessors.

      That is how you get Webview2 all over the place.

      • >Why do you think Windows is currently such a mess?

        There have been some writings and posts here about Microsoft. Here is one from last spring, from a guy that was long time Windows core developer and moved to Azure group. It's well worth reading, what he writes about challenges they have had and most likely still have if not even worse now.

        https://isolveproblems.substack.com/p/how-microsoft-vaporize...

        and the related HN thread

        https://news.ycombinator.com/item?id=47616242

        And from what I've understood old chaps like Dave Cutler are involved much less than they were for a very long time.

        • I have read that, yeah it also shows a lot how things changed.

          As per his interview on Dave's Garage, besides being of an age where he really doesn't need to work, at the time he was involved on getting Linux running on XBox on Azure, apparently Microsoft uses idle consoles from XBox Cloud for AI.

          https://youtu.be/xi1Lq79mLeE?t=10743

      • i fear the times when the chip makers change their architectures and the OS makers have to change the inner working of their OSes but then again, I guess there is an equivalent in the chip-making industry as well
    • > One thing that scares me a little is whether there are younger developers, say, 25-40, who can and want to pick up the mantle of Windows internals gurus.

      A similar "brain drain" has occurred in macOS (formerly known as OS-X) over the years, as evident in man page documentation for "newer" daemons shipped. An easy way to verify this is to run:

        ps -A | awk '{ print $4 }' | grep 'libexec/.*[a-z]d$'
      
      And compare the man pages for the daemons running with the man page for `launchd`.

      While this exercise is illuminating, it is also depressing IMHO.

      • > formerly known as OS-X

        It was never OS-X, it was OS X, and originally Mac OS X, as in the one after Mac OS 9. The Mac prefix was dropped with Lion (10.7). The Mac OS lingo having itself been introduced with 7.6, before that the OS core was called System.

      • And the documentation, now it is mostly generated, the famous Apple books are now gone, at least the archive is still available.
    • As someone within that age group who moved back to Windows for development and entertainment, I find systems programming on Windows more fun and engaging than on competition OSs.

      Oddly enough the open-source nature of the latter kind of takes away some of the thrill. Eerything is just... there, whereas with Windows there's always quite a bit of digging and investigation involved. Or maybe this is Stockholm syndrome; I dunno.

      • Funny feelings in my tummy reading this! I worked with winapi back in late 2000s to early 2010s and I remember having some great fun with it. Although there was MFC and WPF I didn't want to learn them because I wanted the fastest & leanest (catch the reference :) executable I could get; I'd then run gnu strip over the .exe too.

        Stack Overflow was essential to figure out arcane flags that could solve my issues (when even MSDN, another great site with its examples, couldn't) and Raymond was very present in SO at that time, iirc he replied to one of my questions too. That's when I found his blog, always great reads!

        Now I've been a 14-year Linux user and none of the toolkits and libraries give anything close to the winapi experience.

        • > fastest & leanest

          Surely you mean leanest and meanest :P

          I'd also say that tooling on Windows is simultaneously better and easier to use than on Linux; the noob case of green play button in an IDE is taken care of, but if you want detailed performance and memory profiling, record-replay debugging, hot-reload, all of this is straightforwardly available on Windows.

    • I was a dev on the Visual Studio and Windows teams in the 90s. I’m retired but mentor CS students at two local universities.

      I haven’t had a student in two years that was even remotely interested in ring-0, internals, or really understanding a debugger.

      I’m not being critical; they are just focused on higher level abstractions.

      • > I’m retired but mentor CS students at two local universities.

        > I haven’t had a student in two years that was even remotely interested in ring-0, internals, or really understanding a debugger.

        I know quite a lot of such people (even in student age) who are interested in such topics. I really have a feeling that you chose the wrong students at the wrong universities.

        Evidence for my point: rather recently, No Starch Press published quite a lot of about such topics - I am rather certain that a publisher knows quite well which kinds of books do or don't sell well at a given time:

        - The Book of Debugging https://nostarch.com/book-of-debugging

        - The Linux Memory Manager https://nostarch.com/linux-memory-manager

        - The Art of 64-Bit Assembly, Volume 2 https://nostarch.com/art-64-bit-assembly-v2

        - The Ghidra Book, 2nd Edition https://nostarch.com/ghidra-book-2e

        - Building a Debugger https://nostarch.com/building-a-debugger

        - Microcontroller Exploits https://nostarch.com/microcontroller-exploits

        - System Programming in Linux https://nostarch.com/system-programming-linux

        - The Art of ARM Assembly, Volume 1 https://nostarch.com/art-arm-assembly-volume-1

        - Getting Started with FPGAs https://nostarch.com/gettingstartedwithfpgas

        - The Book of I²C https://nostarch.com/book-i%C2%B2c

        • > System Programming in Linux https://nostarch.com/system-programming-linux

          Any ideas how this compares with Kerrisk's The Linux Programming Interface? I've only so much time to read one 1000+ page book...

          • > Any ideas how this compares with Kerrisk's The Linux Programming Interface?

            Unluckily, I don't know, but comparing the Table of Contents for both books

            > https://man7.org/tlpi/toc-short.html

            > https://nostarch.com/system-programming-linux

            I would claim that The Linux Programming Interface covers a broader range of topics, and I also think this book goes more in depth. On the other hand, System Programming in Linux seems to be more pedagogical, and is more targeted towards people who profit from doing exercises and programming projects to get their hands dirty.

        • And for some books about Windows system programming:

          - Windows Internals, Parts I and II

          - Windows 10 System Programming, Parts I and II

          - Windows Kernel Programming, Second Edition

          - Programming Windows, 5th and 6th Editions

        • > you chose the wrong students at the wrong universities

          I work with Duke University, the University of North Carolina, and Carnegie Mellon.

          You don't know that "young engineers" are buying those books. I didn't claim that no one is interested in low-level development. My point is that most younger developers couldn't explain the difference between a mutex and a critical section, or how the OS handles a thread quantum, if their lives depended on it.

          I could list a dozen new books on how to build an LLM from scratch. That doesn't mean that most developers understand LLM internals.

          As an aside, I love your username. I have a tattoo of Aleph One. ;-)

          • > My point is that most younger developers couldn't explain the difference between a mutex and a critical section, or how the OS handles a thread quantum, if their lives depended on it.

            To my knowledge this is taught in some "Operating System" course, and typically students have to do a hands-on implementation of at least some central parts of an operating system. So I guess these students simply did not pay attention in the respective course. :-(

            • I'll help you out: all CS students are required to take an entry-level OS course. Again, you are making incredibly weak arguments given that most students are required to take an English writing course but then can't remember most of it.

              It isn't that they didn't learn it: the issue is that most CS students graduate and work in areas that require zero OS knowledge. For example, when would I spin up a thread versus a fiber? Even ring-3 devs need to have some level of understanding if they want to create performant software.

              First, I didn't work with the right universities. Now the students I work with "didn't pay attention".

              /ignored

      • It’s fun to understand just for intellectual curiosity’s sake but the number of people who get to work on shipping code where ring-0 knowledge is useful has to be minuscule as a percentage. It feels like a minor miracle I got to work on device drivers in my career. I imagine students might be more worried about completing assignments and good grades than the ins and outs of debuggers, too, though understanding those is important for problems one is more likely to come across in work projects versus smaller school assignments.
        • > the number of people who get to work on shipping code where ring-0 knowledge is useful has to be minuscule as a percentage

          I wonder if it it is more that the percentage of people who choose to dedicate themselves to that type of work is miniscule. I work in graphics and performance, and it seems similar.

          Few people really work on it specifically at any given company, and I've heard people warn others that there are few jobs in it.

          But video game companies really want people for those roles and will pay well because they're hard to find. Still, few programmers show any interest in specializing in those skills. If you're passionate about it and willing to learn the details you'll eventually find a lot of job opportunities. I know people who want to work with this and do so - I know many more who have specifically said they want to stay away from it. I don't know anyone who wants to and can't.

          • How would one start learning about it?
            • Before listing resources, let me state that the single most important thing is to learn by doing. Write a lot of rendering code, experiment a lot with your own ideas and variations on each exercise. Once you move past the basics don't be afraid to spend several evenings in a row failing to fix what in hindsight seems like an embarrassingly simple math error. That is how you go from having read and sort of understood something to really knowing it. Also, the best approach is usually to devour as much content as you can from several sources to get a wider perspective.

              There are a lot of great resources out there. The best modern beginner friendly resource I am aware of is Scratchapixel[0]. Back when I was first learning 3D I used to follow tutorials on places like NeHe Productions[1], which is probably a bit dated these days.

              For more comprehensive information on all kinds of techniques, with examples from big games for each technique, the absolutely best resource is the book Real-Time Rendering[2].

              If you're interested in ray-tracing rather than rasterization (i.e. more film than video games) a lot of people recommend "Ray Tracing in One Weekend"[3]. If you want to learn state of the art ray-tracing in depth, with all the math and and theory, the best resource is "Physically Based Rendering: From Theory to Implementation"[4], which is freely available online.

              [0] https://www.scratchapixel.com/

              [1] https://nehe.gamedev.net/

              [2] https://www.realtimerendering.com/

              [3] https://raytracing.github.io/

              [4] https://www.pbrt.org/

            • If you want to learn more about how computers work I would recommend the following:

              - The OS Dev wiki - Open Source Firmware Conference — TKey (shameless plug) - Tiny Tapeout - wafer.space

              In rough hierarchical order from software to metal.

      • Arguably the lower level abstractions are more interesting too! how exactly Windows does ring-0 is less interesting than writing your own ring-0! And unless you care about writing driver-level software for Windows or contributing to the kernel, learning this is also less useful.

        I'm essentially arguing that unless you work at MSFT, there's next to no reason to learn that specific abstraction layer.

        • > I'm essentially arguing that unless you work at MSFT, there's next to no reason to learn that specific abstraction layer.

          Really? Understanding the cost of ring transitions is incredibly useful. I recently consulted with a company that was having horrible performance issues, and it came down to the fact that the primary developer didn't know that certain Win32 calls forced ring transitions. The entire fix was switching from a mutex to a critical section (one causes a ring transition, the other doesn't).

          Treating the OS like an impenetrable black box will bite upcoming engineers/companies... eventually.

      • Speaking of debuggers...

        I still twitch whenever someone says "use ddd" and they are not referring to Evans' seminal work.

        :-D

    • > who can

      Probably enough to keep Windows going, at least.

      > and want to

      Not if the pay or location is uncompetitive.

    • Why do you think that is not already happening within Microsoft?
    • One trend to watch is AI cheat devices. Instead of running detectable software they have a fully separate device that uses AI for object detection and aimbotting. If cheaters move to using those, then the argument for kernel mode anticheat weakens. And that is the cornerstone keeping gamers on windows.
      • Now I'm looking forward to these AI cheat devices. It's going to be hilarious if the kernel anticheat malware finally gets killed by AI aimbots of all things.
        • The term "AI" used in video games is not really related to the current LLM craze.
      • While I’m not sure if this is a bot (where did vidya enter the convo?), game hacking on both cheat and anticheat side has genuinely deep Windows internals knowledge (admittedly somewhat lopsided, but deep nonetheless)
        • Producer and consumer sides are tied. If there is strong demand for windows development then there will be money sloshing around which will attract devs. I predict demand will decrease, at least for this particular niche.

          To go even further off topic, being called a bot is certainly a wake up call for me that the internet is dying, and I'm not ready for it, and need to reposition myself asap somehow.

          • Windows demand is overwhelmingly corporate. Just like Nvidia is barely bothering with gamers at the moment, Windows is not being kept alive for games.
    • Obviously it's Copilot /s
    • Maybe they'll just train Copilot on their code to help them.
  • Only passingly related, some fun rust stack-allocation insanity by my 17yo son:

    https://ogghostjelly.github.io/slog/alloca.html

    • This is great. I'm learning Rust myself and your son's article contributed to my knowledge.

      I'm also very impressed by part 2. I have my own lisp but I haven't managed to implement a compiler or code generation yet. Really enjoyed reading about the hashmap too. The textbook solution to collisions is probing and comparison. It never occurred to me that I could just resize the underlying array until the collisions disappear altogether.

    • 17? You should be very proud. This is good work for anyone, but especially at his age!
    • Very impressive for a 17yo!
  • Related to the above, two important concepts to know w.r.t a stack are "Red Zone" and "Guard Pages".

    Raymond Chen again;

    Why do we even need to define a red zone? Can’t I just use my stack for anything? - https://devblogs.microsoft.com/oldnewthing/20190111-00/?p=10...

    A closer look at the stack guard page - https://devblogs.microsoft.com/oldnewthing/20220203-00/?p=10...

    • I wonder how Linux manages without explicit _chkstk? In my experience, it feels like MAP_GROWSDOWN regions have way more than 1 guard page below its start — I can poke like a megabyte lower than its start, and the kernel will grow the memory region into there just fine.
      • It doesn't. This causes the StackClash vulnerability.
      • Preventing stack guard-page hopping - https://lwn.net/Articles/725832/

        According to this article, allocation in page sizes with implicit probing is used;

        Stack clash mitigation in GCC, Part 3 (-fstack-clash-protection option) - https://developers.redhat.com/blog/2020/05/22/stack-clash-mi...

        • Thank you for the context. But still.

          You have a desirable performance optimization feature —used in every Linux program— that happens to interfere with a lousy exploit mitigation.

          No one should ever need more than 64kBs for a stack anyways.

          • > No one should ever need more than 64kBs for a stack anyways.

            Well, if people would stop storing anything except than return addresses on the stack, yeah, probably even 32 KiB of stack would be enough for anyone. It'd also single-handedly stop all kinds of stack-smashing attacks, too: can't overwrite a return address on the stack if nothing stores data on the stack except the CALL/RET instructions.

            Unfortunately, the current zeitgeist is still to have "writeable stacks" which are only moderately less horrible for the security than "executable stacks".

  • [flagged]