
[CIC 코드 읽기] 공식 구현의 Tensor와 Gradient 흐름
공식 rll-research/cic 구현에서 64D skill sampling, CPC tensor, k-NN intrinsic reward, gradient 경로와 DDPG update가 실제로 어떻게 연결되는지 추적한다.

공식 rll-research/cic 구현에서 64D skill sampling, CPC tensor, k-NN intrinsic reward, gradient 경로와 DDPG update가 실제로 어떻게 연결되는지 추적한다.

DIAYN과 DADS 이후 CIC가 transition-skill contrastive learning과 particle entropy를 결합해 고차원 연속 skill을 학습하는 원리, 실제 DDPG 학습 구조, URLB 실험과 한계를 정리한다.

DIAYN의 상태 다양성에서 출발해 DADS의 조건부 mutual information, skill dynamics, intrinsic reward, SAC 학습과 latent-space planning을 비교 중심으로 정리한다.

DIAYN-PyTorch의 transition 흐름을 추적하고 robot_obs, behavior_features, skill horizon, safety constraint를 갖춘 범용 로봇 구조로 연결한다.

DIAYN 논문의 정보이론 목적함수부터 discriminator 보상, SAC 구현, VIC와의 차이, 실험 결과와 로봇 적용 시 한계까지 정리한다.

3Blue1Brown Deep Learning Chapter 7을 바탕으로 Transformer MLP의 up projection, activation, down projection과 residual update를 살펴보고 feature, neuron, superposition 관점에서 지식 표현을 정리한다.

3Blue1Brown Deep Learning Chapter 6를 바탕으로 query, key, value, scaled dot-product, causal mask, softmax, weighted sum과 multi-head attention의 계산 순서와 행렬 차원을 정리한다.

3Blue1Brown Deep Learning Chapter 5를 바탕으로 입력 문장이 token, embedding, attention, MLP, unembedding과 softmax를 거쳐 다음 token의 확률분포가 되는 전체 흐름을 정리한다.

3Blue1Brown Deep Learning Chapter 3과 4를 바탕으로 출력 오차가 뒤쪽 layer부터 전달되며 각 weight와 bias의 gradient를 계산하는 과정을 직관과 수식으로 정리한다.

3Blue1Brown Deep Learning Chapter 2를 바탕으로 cost function, gradient, negative gradient와 learning rate를 연결해 신경망의 parameter가 어느 방향으로 수정되는지 정리한다.