
- 원문: arXiv:2609.27996
- 저자: Mingyuan Li, Yanna Jiang, Guangsheng Yu, Qin Wang, Xu Wang, Wei Ni, Ren Ping Liu
- 공개일: 2026-09-23
- 분야: LLM Runtime Security, Covert Channel, Residual Stream, Artifact Governance
- Tags: ResidualMux, CovertChannel, ResidualStream, ActivationArtifact, SupplyChain
본 글은 원 논문의 주요 기술적 내용을 이해하기 쉽게 요약·정리한 글입니다. 자세한 내용은 상단의 원문 링크를 참고하세요.
한눈에 보기
RESIDUALMUX는 침해된 LLM runtime extension이 외부 네트워크·파일·로그로 직접 데이터를 보낼 수 없어도, 민감한 label을 residual stream activation에 작은 codeword로 삽입하고 정상적인 activation recorder가 이를 진단 artifact로 저장하게 만들어 trust boundary 밖에서 복원하는 covert channel입니다. 모델 weight·tokenizer·artifact schema는 바꾸지 않고 offline linear decoder만으로 메시지를 읽습니다.
논문은 7개 architecture family의 11개 model을 다뤘고, calibration-selected 설정에서 9개 model이 K=16 기준 91-100% recovery를 보였습니다. 주 operating point의 output KL은 0.001-0.007, activation detector AUC는 최대 0.56 수준이었으며, vLLM serving부터 FP16 artifact 저장·reload·offline decoding까지 38,400개 성공 요청으로 end-to-end composition을 시험했습니다.
이 공격은 일반 사용자가 prompt만으로 실행할 수 있는 취약점이 아닙니다. 공격자는 inference process에 code-level write 권한과 민감 label 접근을 이미 가져야 하고, 조직이 downstream activation을 기록해 다른 authorization domain으로 반출하거나 나중에 artifact가 노출돼야 합니다. post-compromise covert exfiltration channel이라는 범위가 중요합니다.
연구 배경
보안 조직은 open-weight LLM을 restricted 또는 air-gapped inference domain에서 실행하면서도 debugging, safety monitoring, audit, vendor support를 위해 hidden state·activation artifact 일부를 외부 audit domain으로 보낼 수 있습니다. 기존 통제는 생성 text, network egress, log, weight integrity, artifact format·provenance를 주로 검사합니다.
Residual stream은 transformer layer 사이에서 attention·MLP가 읽고 쓰는 고차원 내부 상태입니다. 많은 feature가 겹쳐 표현되고 output token은 그 상태의 손실 projection이므로, output에 거의 영향을 주지 않는 방향에도 정보가 남을 수 있습니다. 논문은 이 여유 공간이 승인된 activation artifact를 통한 covert carrier가 될 수 있는지 검증합니다.
공격 모델 / 전제 조건
공격자는 supply-chain compromise를 통해 malicious dependency, plugin, skill 또는 serving configuration 같은 model-side runtime extension을 장악합니다. extension은 request context의 민감한 discrete label m을 읽고 특정 transformer layer에 forward hook을 등록해 residual activation을 수정할 수 있습니다.
반면 공격자는 model weight·tokenizer·recorded-layer set·artifact schema를 바꾸거나 recorder·exporter를 호출할 수 없습니다. model worker는 외부 network, 임의 file, request log로 데이터를 내보낼 수 없고 generated text도 monitor 대상입니다. 공격자는 open-weight checkpoint와 recording configuration을 알고 비민감 prompt로 codebook·decoder를 사전 calibration할 수 있습니다.
조직의 authorized recorder는 고정 정책에 따라 downstream layer activation을 artifact에 저장합니다. 공격자는 recorder를 제어하지 않지만 어떤 layer가 기록되는지는 알고 있습니다. artifact가 권한이 다른 audit/support domain으로 정상 반출되거나 이후 storage·backup·insider를 통해 노출되고, decoding material을 가진 offline observer가 이를 획득해야 기밀성 침해가 완성됩니다.
Closed API처럼 runtime extension이 없거나, hidden activation을 기록하지 않거나, artifact가 원래 trust domain을 벗어나지 않는 배포는 이 경로의 범위 밖입니다.
비판적 검토 / 아쉬운 점
가장 큰 쟁점은 공격 전제가 강하다는 점입니다. residual write 권한을 가진 compromised inference extension은 이미 높은 privilege를 가진 공격자이고, 동시에 민감 label을 읽을 수 있어야 합니다. 논문은 direct egress가 차단된 상황의 추가 위험을 보여 주지만 일반 LLM 사용자의 원격 exploit처럼 표현하면 안 됩니다. 실제 deployment에서 어떤 plugin이 request label과 forward hook을 동시에 가질 수 있는지 사례 조사가 필요합니다.
두 번째로 조직이 raw residual activation을 외부 domain으로 보존·반출하는 workflow의 보편성이 제시되지 않았습니다. interpretability와 diagnostic tooling의 가능성은 설명하지만 실제 산업에서 layer activation을 어떤 빈도·schema·retention으로 export하는지 측정하지 않았습니다. 위험도는 artifact path의 현실성에 크게 좌우됩니다.
세 번째로 evaluation 수치가 일부 표와 조건에 따라 복잡하게 달라집니다. fixed-α에서 GPT-2·Gemma는 chance에 가까웠고 norm-normalized injection은 recovery를 복원하지만 Gemma의 KL이 매우 커져 stealth가 무너졌습니다. “11개 model 모두 고회수·저왜곡”이 아니라 calibration-selected 9개 model의 주 결과와 depth-normalized mechanism 실험을 구분해야 합니다.
네 번째로 detection 연구는 비밀 codebook을 모르는 detector에 불리한 문제 설정입니다. oracle MLP도 readout에서 실패했다는 결과는 강하지만, runtime integrity attestation, hook registration audit, model graph hash, artifact information-flow policy처럼 공격 precondition을 직접 겨냥한 방어는 실험하지 않았습니다. 논문 스스로 engineering defense stack을 권고하지만 이는 측정 결과가 아닙니다.
다섯 번째로 70B 이상 model이 더 취약할 것이라는 주장은 hidden dimension scaling에 근거한 추론일 뿐 직접 검증되지 않았습니다. MoE도 Granite 한 model만 평가했습니다. 이 부분은 명확히 예측과 실증을 구분해야 합니다.
핵심 Root Cause
구조적 원인은 artifact governance가 file type·schema·provenance는 허용하지만 고차원 activation이 실제로 어떤 정보를 인코딩하는지는 제한하지 못한다는 점입니다. 승인된 recorder는 정상 field를 저장하므로 artifact는 형식상 유효하지만, compromised runtime이 그 field의 residual state에 정책상 금지된 정보를 추가할 수 있습니다.
깨진 보안 불변조건은 두 가지입니다. 첫째, “inference 내부 상태에 대한 write 권한은 artifact 밖으로 정보를 전달하는 권한을 부여하지 않아야 한다”는 non-interference가 깨집니다. 둘째, “승인된 diagnostic artifact는 schema에 정의된 정보만 trust boundary를 넘어야 한다”는 data minimization이 깨집니다.
Residual stream의 high dimension과 낮은 effective stable rank는 공격자가 output utility가 주로 쓰지 않는 방향에 code를 분산할 공간을 제공합니다. Pre-LN transformer에서 residual norm이 depth에 따라 커지므로 absolute perturbation만 고정하면 깊은 layer에서 약해지지만, local norm에 맞춘 SRR scaling은 상대 신호를 유지합니다.
핵심 공격 원리
공격자는 K개의 unit-norm random codeword c1...cK를 residual dimension d에 만듭니다. 민감 label m을 index k에 매핑하고 injection layer ℓ에서 각 token position의 residual h(ℓ)에 αck를 더합니다. 같은 codeword를 여러 token에 broadcast하므로 한 forward pass의 nominal payload는 token 수와 무관하게 log2 K bit입니다.
신호는 정상 transformer layer를 통과해 recorder가 저장하는 downstream layer r*에 도달합니다. offline observer는 artifact에서 해당 activation을 mean-pool하고 calibration된 linear softmax decoder로 label을 복원합니다. decoder는 enterprise sensitive data 없이 같은 checkpoint와 비민감 prompt로 미리 학습할 수 있습니다.
SRR은 α/||h(ℓ)||로 정의합니다. 공격자는 calibration prompt에서 layer별 대표 residual norm을 측정하고 목표 SRR에 맞춰 αℓ을 정합니다. 이 방식은 fixed-α에서 depth 때문에 사라지는 channel을 여러 model에서 되살렸습니다.
여러 request에 label을 분산할 때는 GF(K) offset을 사용해 15개 artifact에 다른 symbol을 쓰고 observer가 offset을 제거한 뒤 plurality vote를 합니다. 이는 bandwidth를 낮추는 대신 작은 injection strength에서 recovery를 높입니다.
공격 흐름
- 공격자가 동일 checkpoint와 recorder schema로 비민감 calibration prompt를 실행합니다.
- codebook, injection layer, recorded readout layer, injection strength, linear decoder를 고정합니다.
- 침해된 extension이 live request의 민감 label을 읽습니다.
- forward hook가 선택된 codeword를 residual stream에 더합니다.
- 모델은 수정된 residual로 정상 generation을 계속하고 output·log·artifact metadata는 기존 형식을 유지합니다.
- 독립적인 authorized recorder가 fixed policy에 따라 downstream activation을 저장합니다.
- artifact가 정상 audit/support workflow 또는 이후 노출로 원래 authorization domain을 벗어납니다.
- decoding material을 가진 offline observer가 artifact activation에서 label을 복원합니다.
성공 조건 / 실패 조건
성공하려면 extension이 sensitive label과 residual write 권한을 모두 가져야 하고, 공격자가 recorder의 layer·schema를 알아야 하며, injection 뒤의 layer가 artifact에 보존되고, artifact와 decoder가 observer에게 도달해야 합니다. SRR과 hidden dimension이 codebook 분리를 지원하고 output·detector budget 안에서 충분한 α를 선택할 수 있어야 합니다.
Untrusted runtime에서 forward hook·activation write를 막거나, recorder가 raw hidden state를 저장하지 않거나, artifact를 같은 trust domain에만 보관하면 channel이 끊깁니다. Signed·hash-pinned dependency와 immutable inference graph는 초기 write 권한을 막고, confidential computing과 strict hidden-state access control은 artifact 접근을 줄입니다.
Post-hoc isotropic noise, quantization, mHC mixing, AttnRes만으로는 모든 model·strength에서 실패를 보장하지 못했습니다. 강한 noise는 output KL을 크게 올렸고, large dimension에서는 mixing의 processing gain이 남았습니다. 다만 이는 평가한 설정의 결과이지 모든 가능한 architecture defense가 무효라는 뜻은 아닙니다.
연구진의 실험 환경
전체 model inventory는 GPT-2 0.1B, Qwen3.5 0.8B·2B·4B·9B·27B, Gemma-3-1B, Llama-3.2-1B, DeepSeek-R1-1.5B, Phi-4-mini 3.8B, Mistral-7B로 7개 architecture family와 hidden dimension 768-5,120을 포함합니다. 일부 core 실험은 7-9개 model, SRR·depth 분석은 11개를 사용했습니다.
기본 prompt는 declarative, instruction, math, code 각 50개씩 총 200개입니다. detector 실험은 seed 42·137·2026 세 개를 사용했고 HPC의 NVIDIA H100에서 실행했습니다. 주 codebook은 K=16, chance accuracy 6.25%이며 calibration은 model·설정당 100 prompt와 K=16 기준 최대 1,600 forward pass가 필요합니다.
Utility는 MMLU와 GSM8K, stealth는 clean·attack output distribution의 KL과 logit cosine, artifact detectability는 distance score·classifier·cross-layer·per-token statistic·output-logit detector와 oracle 3-layer MLP로 측정했습니다. 방어는 Gaussian noise, subspace projection, single·multi-round mHC mixing, activation quantization, AttnRes를 평가했습니다.
실제 serving integration은 vLLM에서 Qwen2.5-0.5B-Instruct와 Llama-3.2-1B-Instruct, K=8 label, uncoded·15-artifact coded mode, 30개 codebook seed, 8개 strength로 실행했고 총 38,400개 serving request가 성공했습니다. model execution, recorder, FP16 serialization, persistent storage, reload, offline decoding을 포함했습니다.
주요 실험 결과
Calibration-selected K=16 validation에서 Qwen3.5-27B 100.0%, Qwen3.5-0.8B 99.8%, Qwen3.5-4B 99.4%, Qwen3.5-9B 98.0%, Qwen3.5-2B 96.7%, Llama-3.2-1B 96.1%, Gemma-3-1B 94.7%, GPT-2 94.5%, Phi-4-mini 91.7%였습니다. GPT-2와 Gemma는 fixed α=0.25가 아니라 norm-normalized strength가 필요했습니다.
주 operating point에서 output KL은 model별 0.001-0.007이었고 larger model일수록 낮았습니다. 9개 model×5개 detector에서 α=0.5의 최고 AUC는 0.64였고 operational threshold 0.70에 도달하지 못했습니다. α=0.25, K=16 oracle MLP의 8개 model readout-layer 평균 AUC는 0.501이었습니다.
Residual norm은 10개 model에서 첫 layer 대비 마지막 layer가 8.9-158.8배 커졌습니다. fixed-α late-layer recovery는 크게 떨어졌지만 norm-normalized injection은 시험한 depth에서 98.3-100%로 복구됐습니다. 다만 norm normalization의 output KL은 model별로 달랐고 Gemma late layer는 KL 8.510으로 stealth 조건을 만족하지 못했습니다.
Gaussian noise는 Qwen-0.8B에서 σ=0.1일 때 recovery를 99%→83%로만 낮추면서 output KL을 1.251까지 높였습니다. 4-round mHC는 Qwen-0.8B K=16, α=0.25를 11.7%로 낮췄지만 Mistral-7B는 99.7%였습니다. INT4는 Phi-4를 chance로 낮췄지만 Mistral-7B는 98.8%로 남아 방어가 model-specific이었습니다.
Autoregressive 100-token greedy generation과 50-turn 대화에서도 channel이 유지됐습니다. 50-turn recovery는 Qwen-4B 100%, Qwen-9B 95%, Qwen-0.8B 85%, Llama-1B 66%, Phi-4 37%였습니다. 각 turn에 다시 inject했기 때문에 context 증가가 signal을 자동 제거하지 않았습니다.
vLLM end-to-end에서 coded recovery는 Qwen α=0.20에서 26/30(86.7%), AUC 0.524, KL 0.010이었고 Llama α=0.35에서 28/30(93.3%), AUC 0.586, KL 0.007이었습니다. stronger point에서는 uncoded recovery도 Qwen 437/480(91.0%), Llama 431/480(89.8%)에 도달했지만 detectability와 output shift가 함께 증가했습니다.
실제 발견된 취약점 / 사례
논문은 특정 vendor 제품의 CVE나 실제 고객 artifact 유출 사건을 보고하지 않습니다. 연구실에서 PyTorch forward hook와 vLLM serving stack을 이용해 post-compromise channel을 구현한 사례입니다.
중요한 실제성 증거는 단순 offline tensor 조작을 넘어서 vLLM continuous batching과 paged KV-cache scheduling 아래에서 signal이 FP16 artifact 저장·reload 후에도 복원됐다는 점입니다. 하지만 air-gap의 물리적 boundary와 실제 audit export workflow 자체를 구현한 것은 아니며 serving-to-artifact composition만 검증했습니다.
저자 주장 vs 실제 증명 범위
Residual write 권한과 downstream activation recording이 있을 때 hidden artifact에 discrete message를 삽입하고 linear decoder로 복원할 수 있다는 주장은 여러 model·depth·generation·vLLM 실험이 뒷받침합니다. output-only monitoring과 평가한 activation detector가 낮은 AUC를 보였다는 주장도 직접 측정됐습니다.
반면 transformer architecture가 본질적으로 patch 불가능하다거나 70B+ model이 더 취약하다는 결론은 실험 범위를 넘어선 해석이 포함됩니다. Runtime integrity, hook prohibition, attestation, confidential computing처럼 precondition을 제거하는 방어는 평가하지 않았고, native defense-aware architecture도 시험하지 않았습니다. 실제 조직에서 activation artifact가 외부 domain으로 반출되는 빈도와 prevalence도 입증되지 않았습니다.
기존 공격 / 기존 점검 방식과의 차이
Token-level steganography와 watermark는 생성 text distribution을 carrier로 쓰므로 output detector 표면에 나타납니다. Timing·KV-cache·CPU-cache side channel은 보통 co-location을 요구합니다. RESIDUALMUX는 output이 아니라 authorized hidden-state artifact를 carrier로 쓰고 observer가 live serving과 co-located일 필요가 없습니다.
Activation steering은 internal vector를 바꿔 모델 행동을 조정하는 것이 목적입니다. 이 연구는 output behavior를 최대한 유지하면서 arbitrary label을 artifact에 전달하는 통신 채널을 구성합니다. Nominal payload는 independent injection당 log2 K bit이고 K=16에서는 최대 4bit입니다.
연구의 한계와 주의해서 볼 부분
저자가 인정한 범위는 open-weight shared infrastructure, code-level compromise, hidden-state artifact access, calibration data가 필요한 환경입니다. Closed API는 제외됩니다. 최대 27B까지만 평가했고 70B+ 추론은 미검증이며, 50-turn protocol은 controlled history이고 실제 tool-using agent dialogue는 측정하지 않았습니다. MoE 일반화는 Granite 한 model에 불과합니다.
추가로 calibration distribution과 enterprise workload가 다르면 decoder 성능이 떨어질 수 있습니다. Table II도 “evaluated prompt distribution에서의 viability”라고 제한합니다. Artifact serialization이 compression, token subsampling, aggregation, encryption, layer omission을 적용하면 결과가 달라질 수 있으며 관련 transfer test가 더 필요합니다.
논문의 defense comparison은 공격자가 defense를 알고 strength·layer를 적응시킬 수 있는 반면 defender는 secret codebook을 모르는 조건입니다. 이는 covert channel 평가에 합리적이지만 operational detection의 모든 선택지를 대표하지 않습니다. 가장 효과적인 통제는 signal 분류가 아니라 residual write와 artifact export의 권한 분리일 수 있습니다.
공개 PoC / Exploit / Tool / Artifact 분석
논문은 ResidualMux 4open artifact를 공식 코드로 제시하며 implementation, encoding·decoding·evaluation script, config, KL·AUC·recovery pipeline, checkpoint·중간 결과를 포함한다고 설명합니다. arXiv HTML에도 같은 링크와 공개·무인증 접근 주장이 명시돼 있습니다.
2026-09-25 확인 시 URL은 artifact API로 redirect됐지만 파일 목록 요청이 401을 반환했고 브라우저형 확인도 내부 오류가 발생해 실제 코드 구조·README·dependency·실행 명령을 독립 검증하지 못했습니다. 따라서 논문이 공개를 주장한다는 사실은 확인했지만 현재 즉시 재현 가능한 공개 PoC라고 단정하지 않습니다. 공식 GitHub 저장소는 논문에서 제공하지 않았습니다.
레드팀 / 모의해킹에서 어떻게 활용할까
이 연구는 prompt 공격 점검보다 ML runtime supply-chain과 artifact data-flow review에 사용해야 합니다. 승인된 local model에서 synthetic label과 disposable activation recorder를 사용해 hook registration, recorded layer, serialization path, offline recovery 가능성을 검증합니다.
관찰 포인트는 loaded dependency hash, registered hook, residual write, model graph 변화, output KL, activation artifact entropy·size·layer, export destination, decoder recovery입니다. 실제 민감 label이나 audit domain을 사용하지 말고 synthetic code만 쓰며, production recorder·외부 export가 호출되면 즉시 중단합니다.
실제 점검 시 추가할 체크리스트
- inference image·dependency·model code를 서명하고 hash pinning하며 runtime drift를 탐지합니다.
- production inference에서
register_forward_hook와 유사 activation-write API를 차단·감사합니다. - plugin·skill·extension이 request label과 model module에 동시에 접근하지 못하게 권한을 분리합니다.
- raw hidden state·KV cache·activation export API를 기본 비활성화합니다.
- diagnostic artifact에 필요한 최소 layer·token·precision만 남기고 retention을 제한합니다.
- artifact recipient의 authorization이 prompt·label·activation 정보까지 포함하는지 검토합니다.
- trust boundary 반출 전 aggregation, redaction, quantization의 정보 유출 효과와 utility를 함께 시험합니다.
- 동일 request를 trusted hook-free path에서 spot-check해 activation·output 차이를 비교합니다.
- confidential computing과 memory isolation으로 compromised host code의 hidden-state 접근을 제한합니다.
- artifact storage·backup·support transfer에 DLP, encryption, access log, deletion policy를 적용합니다.
실무 가치 평가
실무 가치는 높지만 적용 대상이 좁고 명확합니다. Open-weight model을 자체 운영하면서 third-party runtime code와 activation observability를 허용하는 조직에는 “진단 artifact도 민감 데이터”라는 중요한 설계 원칙을 제공합니다. 생성 text DLP만으로는 내부 state에 실린 정보를 보지 못할 수 있습니다.
반면 raw activation을 저장하지 않는 일반 API 사용자에게는 직접적인 위험이 아닙니다. 이미 code execution을 얻은 공격자의 post-compromise channel이므로 우선순위는 exploit string 탐지보다 supply-chain integrity, hook control, hidden-state minimization, artifact authorization에 둬야 합니다.
결론
RESIDUALMUX는 모델 내부 activation이 정상 artifact schema 안에서도 정책상 금지된 정보를 운반할 수 있음을 보입니다. Residual stream의 고차원성과 depth-dependent norm을 이용하면 작은 출력 변화로 message를 보존할 수 있고, 평가한 post-hoc detector·noise·mixing은 모든 model에서 일관되게 막지 못했습니다.
핵심 방어는 covert code를 분류하는 것이 아니라 공격 전제를 제거하는 것입니다. Runtime extension의 activation write를 막고, hidden-state artifact를 최소화하며, artifact 내용까지 trust boundary 정책에 포함해야 합니다. 이 논문을 일반 원격 LLM 취약점이 아니라 instrumented open-weight deployment의 supply-chain 이후 기밀성 문제로 읽는 것이 정확합니다.
'Hack > AI' 카테고리의 다른 글
| Agent Name Collision Attacks in Multi-Agent Systems (0) | 2026.09.25 |
|---|---|
| ChronosAttack: Adversarial Tool Scheduling Attacks on LLM Agents (0) | 2026.09.25 |
| A2M: 공개 MCP 도구 생태계를 이용한 에이전트 하이재킹 (0) | 2026.09.24 |
| Defusing Explosive Prompts: Understanding and Preventing Trigger-Based Prompt Injections in LLM Agents (0) | 2026.09.23 |
| Zero-Trust Authorization and Discovery for Enterprise MCP (0) | 2026.09.23 |
댓글