Title is a bit misleading, and I think intentionally, which kinda irks me. It isn’t a 400x speedup of RAM by means of compression. If I understood correctly, it’s 400x speedup compared to disk read for swap.
Please correct me if I’m wrong here. I’d love some positive news from the IT world that isn’t depressive. (edit: I implied this wasn’t. It is. I’m just not sure of the actual impact)
It’s a little better than your read. The ~400x speedup is in comparison to using ZRAM, which, in very brief terms, compresses memory by creating an in-memory compressed block device and assigning that as your swap space. So your “swap” is actually a compressed chunk of RAM, not on disk.
This was significantly slower than normal memory access because page faulting when looking up something in memory then fetching from swap was, itself, expensive. Regardless of how fast that swap was. When your swap is on a storage device that overhead is comparatively tiny, but when it’s just another chunk of memory suddenly it’s what you’re spending most of your time on.
Fuck yea, RAM swap? I bought both my laptop and desktop when RAM was cheap, and got more memory than I needed. I can safely allocate like 12GB to CRAM!
The win from CRAM is that it is not swap, and operates using mostly normal RAM semantics. If it works as advertised you should be able to allocate most (all?!) of your RAM as CRAM.
I would guess this is only useful in servers, as for day to day use you ideally want the lowest latency you can for responsiveness. That said, if there’s a way you can specify which type processes can use, that may be useful in some cases
Honestly though CPU processing is going to add almost no latency at all.
Almost all latency is sending to the RAM itself, so if you can compress it in a CPU cache before sending it, its nearly free RAM space.
Also systems are fucked fast these days you could probably make something responsive in pure python - other than millisecond timing scenarios like medical and audio stuff
I’m not sure if this is misleading at all. If anything, they should have mentioned compared to ZRAM. Title is incomplete, but I don’t think it is intentionally misleading here. It just looks like they expect the reader to know CRAM is a replacement for ZRAM. The new compression method CRAM with over 400x speedup is compared to ZRAM method:
A new compression model, called CRAM, offers a different path to compression that avoids swap entirely by keeping the compressed data in memory, and it offers up to 452x the performance of ZRAM.
Because CRAM is stored in RAM and treated as RAM, with full cacheline/byte access, it can be accessed in a read-only fashion with little delay; just the cost of hardware-offloaded compression. As a result, CRAM “runs at DRAM speed,” as the creator says in the slide above. While the graph already looks impressive, it’s a logarithmic scale; CRAM, in the worst case, is doing 489 million operations per second versus ZRAM’s 1.1 million. It’s barely comparable.
Even when you enable writes, CRAM is still much faster than ZRAM; 5.4x in the worst tested case of 20% writes. That’s a huge drop from the 452x read-only case, but keep your context; a 5.4x speedup is still titanic.
Is writes enabled CRAM / ZRAM common? If so, then the post title is definitely misleading.
Guess read only RAM becomes… ROM? :D I have no clue either. Maybe there are protected areas in the memory no program has write access to, so it is read only from perspective of the application. Searching the web doesn’t help, because every link I clicked just explains the difference between ROM and RAM.
Hmm… when I think about Rust programming (which is true in C too probably), there are two types of locations our variables can assigned to: Stack and Heap. In example if you have a text string as a literal like “Version 1.0”, that string is located in the Stack memory, because it is unchanging. The Heap gets all those content that can vary and arbitrary long, but its slower. So the Stack content is much smaller, faster and basically read only RAM area (if I understand this correctly). Maybe that is it?
I’m drawing on some old memories here, so I could be mistaken, but I don’t think the stack is read only, not in C anyway or in the underlying machine code. If it is faster it has to do with greater overhead needed managing the larger heap and perhaps being more efficient to push and pop with small offsets to a local stack frame vs large absolute addresses.
I don’t mean the stack is read only (edit: yes I meant that in my previous reply, but got confused myself, I actually never thought the entire stack being read only, I was only thinking about those specific variables and literal strings, sorry for confusion), but certain variables holding values that are only used to read and not change. In example you cannot change literals, therefore they are read only values. In example if you have a program that prints “Hello Lemmy”, that string is a literal that cannot be altered, and it is found in the application itself, as part of the binary. That part maybe is marked as read only?
Right, a stack that was itself read-only would be hard to use! :^D
C definitely lets you modify the things in the stack. (Ah the fun of bugs where you accidentally overwrite other shit on the stack!) I don’t know Rust, but it would not surprise me if Rust only had immutable things on the stack.
No, its not only immutable things on the stack. I mean if you include a literal constant string such as “MIT LICENSE”, that is part of the compiled binary file. So it is unchangeable. Because you cannot change a literal, a “1” is always a “1” in the compiled binary file. And those are basically read only by their nature and loaded into the stack. I think or guess in C it is the same. Or any language for that matter.
Let me use my Raspberry Pi 3b as a practical example of what @vithigar@lemmy.ca was explaining. I run PiHole + Unbound on the Raspberry Pi 3b, with 1GB of ram. The main killer for a little Pi 3b is reading and writing to the SD card. We want to avoid that as much as possible. Block lists get updated weekly, so that’s no big deal, but there is still a lot of traffic information that is being kept track of. The solution is to load all of that into zram. As you can imagine, on a slow and resource constrained system like a Pi 3b, any performance enhancement has a huge impact. A 452x performance increase over ZRAM on a system that mostly lives on ZRAM is massive.
I am not entirely sure about this but I imagine that this is also massive news for live operating systems like Tails and Kali Live.
Title is a bit misleading, and I think intentionally, which kinda irks me. It isn’t a 400x speedup of RAM by means of compression. If I understood correctly, it’s 400x speedup compared to disk read for swap.
Please correct me if I’m wrong here. I’d love some positive news from the IT world that isn’t depressive. (edit: I implied this wasn’t. It is. I’m just not sure of the actual impact)
It’s a little better than your read. The ~400x speedup is in comparison to using ZRAM, which, in very brief terms, compresses memory by creating an in-memory compressed block device and assigning that as your swap space. So your “swap” is actually a compressed chunk of RAM, not on disk.
This was significantly slower than normal memory access because page faulting when looking up something in memory then fetching from swap was, itself, expensive. Regardless of how fast that swap was. When your swap is on a storage device that overhead is comparatively tiny, but when it’s just another chunk of memory suddenly it’s what you’re spending most of your time on.
Fuck yea, RAM swap? I bought both my laptop and desktop when RAM was cheap, and got more memory than I needed. I can safely allocate like 12GB to CRAM!
The win from CRAM is that it is not swap, and operates using mostly normal RAM semantics. If it works as advertised you should be able to allocate most (all?!) of your RAM as CRAM.
I would guess this is only useful in servers, as for day to day use you ideally want the lowest latency you can for responsiveness. That said, if there’s a way you can specify which type processes can use, that may be useful in some cases
Honestly though CPU processing is going to add almost no latency at all.
Almost all latency is sending to the RAM itself, so if you can compress it in a CPU cache before sending it, its nearly free RAM space.
Also systems are fucked fast these days you could probably make something responsive in pure python - other than millisecond timing scenarios like medical and audio stuff
Depends on the workload. This probably would not be great for gaming or audio production, but it might be for video editing and local LLMs
And for gluttonous browsers and browser-based apps
I’m not sure if this is misleading at all. If anything, they should have mentioned compared to ZRAM. Title is incomplete, but I don’t think it is intentionally misleading here. It just looks like they expect the reader to know CRAM is a replacement for ZRAM. The new compression method CRAM with over 400x speedup is compared to ZRAM method:
Is writes enabled CRAM / ZRAM common? If so, then the post title is definitely misleading.
Yeah wait how would you have read only RAM?
Loaded on boot for the OS I suppose?
Probably nice for running TV boxes off even less RAM than they’re already starved for lmao
Guess read only RAM becomes… ROM? :D I have no clue either. Maybe there are protected areas in the memory no program has write access to, so it is read only from perspective of the application. Searching the web doesn’t help, because every link I clicked just explains the difference between ROM and RAM.
Hmm… when I think about Rust programming (which is true in C too probably), there are two types of locations our variables can assigned to: Stack and Heap. In example if you have a text string as a literal like “Version 1.0”, that string is located in the Stack memory, because it is unchanging. The Heap gets all those content that can vary and arbitrary long, but its slower. So the Stack content is much smaller, faster and basically read only RAM area (if I understand this correctly). Maybe that is it?
I’m drawing on some old memories here, so I could be mistaken, but I don’t think the stack is read only, not in C anyway or in the underlying machine code. If it is faster it has to do with greater overhead needed managing the larger heap and perhaps being more efficient to push and pop with small offsets to a local stack frame vs large absolute addresses.
I don’t mean the stack is read only (edit: yes I meant that in my previous reply, but got confused myself, I actually never thought the entire stack being read only, I was only thinking about those specific variables and literal strings, sorry for confusion), but certain variables holding values that are only used to read and not change. In example you cannot change literals, therefore they are read only values. In example if you have a program that prints “Hello Lemmy”, that string is a literal that cannot be altered, and it is found in the application itself, as part of the binary. That part maybe is marked as read only?
Right, a stack that was itself read-only would be hard to use! :^D C definitely lets you modify the things in the stack. (Ah the fun of bugs where you accidentally overwrite other shit on the stack!) I don’t know Rust, but it would not surprise me if Rust only had immutable things on the stack.
No, its not only immutable things on the stack. I mean if you include a literal constant string such as “MIT LICENSE”, that is part of the compiled binary file. So it is unchangeable. Because you cannot change a literal, a “1” is always a “1” in the compiled binary file. And those are basically read only by their nature and loaded into the stack. I think or guess in C it is the same. Or any language for that matter.
I started to think I was hallucinating. I guess what I was thinking is this part https://en.wikipedia.org/wiki/Data_segment . There are read-only data segments too.
Let me use my Raspberry Pi 3b as a practical example of what @vithigar@lemmy.ca was explaining. I run PiHole + Unbound on the Raspberry Pi 3b, with 1GB of ram. The main killer for a little Pi 3b is reading and writing to the SD card. We want to avoid that as much as possible. Block lists get updated weekly, so that’s no big deal, but there is still a lot of traffic information that is being kept track of. The solution is to load all of that into zram. As you can imagine, on a slow and resource constrained system like a Pi 3b, any performance enhancement has a huge impact. A 452x performance increase over ZRAM on a system that mostly lives on ZRAM is massive.
I am not entirely sure about this but I imagine that this is also massive news for live operating systems like Tails and Kali Live.