RecodeX 重构消息,研究人员利用苹果的’LLM in a Flash’技术,在48GB MacBook Pro M3 Max上通过量化与流式加载,成功以5.5+ tokens/秒的速度本地运行Qwen3.5-397B-A17B模型。