论文阅读知识库
这里记录小组共同阅读的论文、讨论和可复用的研究想法。
登录后可以创建和修改论文。
当前论文
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval — DAC'26 · 26
- AIMS: Cost-Efficient LLM-Based Agent Deployment in Hybrid Cloud-Edge Environments — Eurosys · 26
- Scaling LLM Test-Time Compute with Mobile NPU on Smartphones — Eurosys26 · 26
- SolidAttention: Low-Latency SSD-based Serving on Memory-Constrained PCs — FAST · 2026 · 待读
- ICML 2026年论文评述合集(低优先级论文) — ICML · 26
- ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling — ICML · 26
- LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents — ICML · 26
- PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Fine-Grained Expert Merging and Bit-packed Inference — ICML 26 · 26
- ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference — ICML · 26
- V-Droid: Advancing Mobile GUI Agent Through Generative Verifiers — MobiCom'26 · 26
- Act Before It's Too Late: Power-Efficient LLM Inference on Mobile Device — Mobisys · 26
- ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference — Mobisys26 · 26
- Benchmarking and Characterization of Large Language Model Inference on Apple Silicon — SIGMETRICS 2026 · 26
- D²MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving — MobiCom · 25
- Modality Plug-and-Play: Runtime Modality Adaptation in LLM-Driven Autonomous Mobile Systems — MobiCom'25 · 25
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — NIPS’22 · 24
- AutoDroid: LLM-powered Task Automation in Android — MobiCom'26 · 24
- MELTing Point: Mobile Evaluation of Language Transformers — MobiCom · 24
新的论文会自动出现在这个列表中。论文文件按会议和会议年份归档,例如 FAST/FAST'26/paper-title.md。