Rendered at 12:14:07 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
sinuhe69 12 hours ago [-]
Did the author really believe the LLM made the compiler by itself and not distilled (copied) from other compilers?
pjmlp 40 minutes ago [-]
Naturally libraries (models) have to come from somewhere.
serbuvlad 5 hours ago [-]
1. I highly suspect a pretrained-only model could not do this, yes.
2. Most problems anyone has are solved problems or variations on solved problems. This type of distillation is extremely useful.
ganelonhb 9 hours ago [-]
All knowledge is distilled from prior knowledge. I don’t think it matters what the author thinks or doesn’t think. The world is different than the one you knew now.
helpprotactiniu 16 hours ago [-]
How did you handle the various optimizations and CPU architectures? compilers are notoriously easy to build but hard to optimize and maintain.
temo_pacheco 16 hours ago [-]
On optimization I deliberately aimed low: a small set of passes, nothing beyond what I could prove. Everything is verified against clang/gcc output on real software rather than against my own expectations, and the compiler compiles itself and reproduces itself bit-for-bit, so it's its own regression test. In practice the generated code holds up fine — DOOM runs smoothly, so do Lua. I chose provable correctness over speed; optimization is the least finished part of the project and there's real work left there.
gomoboo 13 hours ago [-]
I vouched for you here because this compiler is interesting to me and I'd like to see some discussion on it. All your comments are listed as 'dead' though. Likely because they seem to be LLM-generated. Maybe writing by hand or noting you're using an LLM for translation (if you are) would help.
avadodin 12 hours ago [-]
The post probably should have been show HN too being posted by the author himself. I don't know all of the rules but perhaps it would be better for engagement if you reposted the project for him.
I was looking for a timeline given the title. Oldest path seems to be 3 months ago so $300 and 3 months? That could buy you a second-hand graphics card to run the current batch of local models.
temo_pacheco 7 hours ago [-]
I used my $100-per-month Cloud Code subscription. It was enough to build the compiler from scratch. Although the subscription also allowed me to work on other projects simultaneously, the methodology optimizes resource usage—an important aspect to demonstrate as well. That is why I chose to create a C compiler.
avadodin 7 hours ago [-]
Don't get me wrong. I'm here because I'm interested as well.
- Wall clock time is three months(?)
- Cost $300 possibly shared by other projects(Does CC provide credits or a token measure?)
- How much of your own time was spent interacting with Claude and/or thinking in parallel to his work?
- Did the service go down during those three months(including subpar service by a substitute "flash" model)?
temo_pacheco 6 hours ago [-]
Yes, it involved seven weeks of work using Claude Code sessions—from July 25 to September 14, as indicated by the repository commits.
The total cost was simply the recurring monthly subscription fee. I wasn't granted any extra credits by CC; although I occasionally hit my weekly usage limit—partly due to consumption from other projects—I would simply wait a few hours before resuming work.
I spent six to eight hours a day interacting with the session, and whenever I had the time, I dedicated even more hours to it, as my goal was to get the job done. During the final days, much of the time was spent waiting for regression tests to pass so I could proceed with debugging.
The service remained active throughout the entire seven-week period. I consistently used Opus 4.8 to maintain model consistency during development.
avadodin 6 hours ago [-]
That provides a much clearer picture, thank you.
temo_pacheco 6 hours ago [-]
Thank you for your interest in this topic.
temo_pacheco 7 hours ago [-]
I tried to post as "Show HN", but HN restricts Show HN for new accounts.
temo_pacheco 7 hours ago [-]
I am using an LLM solely to help improve my translation. The compiler works, and the code explains it very well. We can discuss the compiler and the method I used to create it with the help of AI.
pjmlp 4 hours ago [-]
The part of it being written in ARM64 Assembly is the most relevant one, yet again making the point that we don't necessary need to generated third generation source code languages as mechanism to produce executables out of AI tooling.
The tools will keep improving, and as for the non-determinism, we already have to deal with it in language runtimes that make use of GC, JIT (+ PGO), and machine learning based compiler passes.
The fifth programming languages generation is coming.
fuhsnn 3 hours ago [-]
> The part of it being written in ARM64 Assembly
There is a C version under literate/verification, for example "main" is [1][2]:
Still don't get the point, it is clear that the compiler was generated in Assembly, and the C part is the second phase of the bootstrapping process.
fuhsnn 1 hours ago [-]
Well, why treat the assembly version as the main implementation when you have a supposedly equivalent C version around? If this is released as a plain C project I would actually be interested to try it out.
pjmlp 42 minutes ago [-]
Because that was the whole goal of the project, and proves a point in AI capabilities.
> kcc is a formal C17 compiler written entirely in ARM64 assembly....
The main implementation is the 5 GL English language in literate programming form, used to drive the AI to produce the ARM64 assembly.
2. Most problems anyone has are solved problems or variations on solved problems. This type of distillation is extremely useful.
I was looking for a timeline given the title. Oldest path seems to be 3 months ago so $300 and 3 months? That could buy you a second-hand graphics card to run the current batch of local models.
- Wall clock time is three months(?)
- Cost $300 possibly shared by other projects(Does CC provide credits or a token measure?)
- How much of your own time was spent interacting with Claude and/or thinking in parallel to his work?
- Did the service go down during those three months(including subpar service by a substitute "flash" model)?
The total cost was simply the recurring monthly subscription fee. I wasn't granted any extra credits by CC; although I occasionally hit my weekly usage limit—partly due to consumption from other projects—I would simply wait a few hours before resuming work.
I spent six to eight hours a day interacting with the session, and whenever I had the time, I dedicated even more hours to it, as my goal was to get the job done. During the final days, much of the time was spent waiting for regression tests to pass so I could proceed with debugging.
The service remained active throughout the entire seven-week period. I consistently used Opus 4.8 to maintain model consistency during development.
The tools will keep improving, and as for the non-determinism, we already have to deal with it in language runtimes that make use of GC, JIT (+ PGO), and machine learning based compiler passes.
The fifth programming languages generation is coming.
There is a C version under literate/verification, for example "main" is [1][2]:
[1] assembly version: ./literate/driver/43-main.weft
[2] C version: ./literate/verification/50-self-hosting-main.weft
https://github.com/LiterateDrivenDevelopment/kcc/blob/main/l...
asm: https://github.com/LiterateDrivenDevelopment/kcc/blob/main/l...
C: https://github.com/LiterateDrivenDevelopment/kcc/blob/main/l...
> kcc is a formal C17 compiler written entirely in ARM64 assembly....
The main implementation is the 5 GL English language in literate programming form, used to drive the AI to produce the ARM64 assembly.