We run LLM inference on-chain (llama cpp canister), so our canisters allocate multi-GiB heaps. After the -deterministic-tracker GuestOS variant rolled out, loading our Qwen3-1.7B model with a 16K context stopped working. We load the model from stable memory (a virtual file in C++ canisters) into…
This forum reads and writes through the Internet Computer from your browser, so the discussion itself needs JavaScript enabled.