Directory listing for /papers/
A_Careful_Examination_of_Large_Behavior_Models_for_Multitask_Dexterous_Manipulation.pdf
A_Comprehensive_Overview_of_Large_Language_Models.pdf
A_Comprehensive_Survey_on_World_Models_for_Embodied_AI.pdf
A_Path_Towards_Autonomous_Machine_Intelligence.pdf
A_Pragmatic_VLA_Foundation_Model.pdf
A_Survey_of_Large_Language_Models.pdf
A_Survey_on_Vision-Language-Action_Models_for_Embodied_AI.pdf
ACT_Action_Chunking_with_Transformers.pdf
Adaptive_Mixtures_of_Local_Experts.pdf
Addressing_Function_Approximation_Error_in_Actor-Critic_Methods.pdf
An_Image_is_Worth_16x16_Words_Transformers_for_Image_Recognition_at_Scale.pdf
ATOMVLA.pdf
Attention_Is_All_You_Need.pdf
Behavior_Generation_with_Latent_Actions.pdf
Behavior_Prompting_Policy_Demonstrations_as_Prompts_for_Manipulation.pdf
Being-H0.7_A_Latent_World-Action_Model_from_Egocentric_Videos.pdf
BERT_Pre-training_of_Deep_Bidirectional_Transformers_for_Language_Understanding.pdf
Chain-of-Thought_Prompting_Elicits_Reasoning_in_Large_Language_Models.pdf
CogACT_A_Foundational_Vision-Language-Action_Model_for_Synergizing_Cognition_and_Action_in_Robotic_Manipulation.pdf
Conservative_Q-Learning_for_Offline_Reinforcement_Learning.pdf
Cosmos_3_Omnimodal_World_Models_for_Physical_AI.pdf
CoT-VLA_Visual_Chain-of-Thought_Reasoning_for_Vision-Language-Action_Models.pdf
DAPO_An_Open-Source_LLM_Reinforcement_Learning_System.pdf
DAWN_World-Action_Interactive_Models.pdf
DECO_Decoupled_Multimodal_Diffusion_Transformer_for_Bimanual_Dexterous_Manipulation_with_a_Plugin_Tactile_Adapter.pdf
Deep_Reinforcement_Learning_for_Robotics_A_Survey_of_Real-World_Successes.pdf
Deep_Residual_Learning_for_Image_Recognition.pdf
DeepSeek-R1_Incentivizing_Reasoning_Capability_in_LLMs_via_Reinforcement_Learning.pdf
DeepSeekMath_Pushing_the_Limits_of_Mathematical_Reasoning_in_Open_Language_Models.pdf
Denoising_Diffusion_Probabilistic_Models.pdf
DexVLA_Vision-Language_Model_with_Plug-In_Diffusion_Expert_for_General_Robot_Control.pdf
Diffusion_Policy_Visuomotor_Policy_Learning_via_Action_Diffusion.pdf
DIGIT_A_Novel_Design_for_a_Low-Cost_Compact_High-Resolution_Tactile_Sensor.pdf
Direct_Preference_Optimization.pdf
DiT4DiT_Jointly_Modeling_Video_Dynamics_and_Actions_for_Generalizable_Robot_Control.pdf
DiT_Scalable_Diffusion_Models_with_Transformers.pdf
DreamGen_Unlocking_Generalization_in_Robot_Learning_through_Video_World_Models.pdf
DreamTacVLA_Learning_to_Feel_the_Future_for_Contact-Rich_Manipulation.pdf
DreamZero_World_Action_Models_are_Zero-shot_Policies.pdf
Dyna-2_A_1-Million-Hour_Scaling_Law_for_World-Action_Models.pdf
Ego-Pi_VLA_Fine-Tuning_for_Ego-Centric_Human_and_Robot_Data.pdf
Ego2Robot_Scalable_Robot_Data_Synthesis_from_Egocentric_Human_Data.pdf
EgoScale_Scaling_Dexterous_Manipulation_with_Diverse_Egocentric_Human_Data.pdf
Emergence_of_Human_to_Robot_Transfer_in_Vision-Language-Action_Models.pdf
Eureka_Human-Level_Reward_Design_via_Coding_Large_Language_Models.pdf
Fast-WAM_Do_World_Action_Models_Need_Test-time_Future_Imagination.pdf
FAST_Efficient_Robot_Action_Tokenization.pdf
FLARE_Robot_Learning_with_Implicit_World_Modeling.pdf
FlexiTac_A_Low-Cost_Open-Source_Scalable_Tactile_Sensing_Solution_for_Robotic_Systems.pdf
Flow_Matching_for_Generative_Modeling.pdf
ForceVLA2_Hybrid_Force_Position_Control.pdf
ForceVLA_Force_aware_MoE_for_Contact_rich_Manipulation.pdf
From_Foundation_to_Application_Improving_VLA_Models_in_Practice.pdf
Galaxea_G0.5_Technical_Report.pdf
GE-Act_2.0_Pretraining_and_Scaling_a_World-Action_Model_for_Robotic_Manipulation.pdf
Gemini_Robotics_1.5_Pushing_the_Frontier_of_Generalist_Robots_with_Advanced_Embodied_Reasoning_Thinking_and_Motion_Transfer.pdf
Gemini_Robotics_2_Safety_Evaluations.pdf
Gemini_Robotics_Bringing_AI_into_the_Physical_World.pdf
Gemini_Robotics_ER_1.6_Model_Card.pdf
Gemini_Robotics_ER_2_Model_Card.pdf
Gemini_Robotics_On-Device_2_Model_Card.pdf
GeomVLA_Unifying_Scene_Motion_and_Action_in_3D.pdf
GPT-4_Technical_Report.pdf
GPT-6_Astra_System_Card_Chinese_Translation_20260903.pdf
GR-2_A_Generative_Video-Language-Action_Model_with_Web-Scale_Knowledge_for_Robot_Manipulation.pdf
GR-3_Technical_Report.pdf
GR-RL_Going_Dexterous_and_Precise_for_Long-Horizon_Robotic_Manipulation.pdf
GR00T_N1_Open_Foundation_Model_for_Generalist_Humanoid_Robots.pdf
Hi_Robot_Open-Ended_Instruction_Following_with_Hierarchical_Vision-Language-Action_Models.pdf
HOST_Robots_Acquire_Manipulation_Skills_in_Seconds_from_a_Single_Human_Video.pdf
Human-level_Control_through_Deep_Reinforcement_Learning.pdf
I-JEPA_Self-Supervised_Learning_from_Images_with_a_Joint-Embedding_Predictive_Architecture.pdf
Igniting_VLMs_Toward_the_Embodied_Space.pdf
ImageNet_Classification_with_Deep_Convolutional_Neural_Networks.pdf
Implicit_Q-Learning.pdf
Improving_Language_Understanding_by_Generative_Pre-Training.pdf
InternVLA-A1.5_Unifying_Understanding_Latent_Foresight_and_Action_for_Compositional_Generalization.pdf
InternVLA-A1_Unifying_Understanding_Generation_and_Action_for_Robotic_Manipulation.pdf
InternVLA-M1_A_Spatially_Guided_Vision-Language-Action_Framework_for_Generalist_Robot_Policy.pdf
JEPA-WAM_Learning_Vision-Language-Action_Policies_with_Joint-Embedding_World_Modeling.pdf
Knowledge_Insulating_Vision-Language-Action_Models_Train_Fast_Run_Fast_Generalize_Better.pdf
Language_Models_are_Few-Shot_Learners.pdf
Language_Models_are_Unsupervised_Multitask_Learners.pdf
Large_Language_Models_A_Survey.pdf
Latent_Action_Pretraining_from_Videos.pdf
Learning_to_Predict_by_the_Methods_of_Temporal_Differences.pdf
Learning_Transferable_Visual_Models_From_Natural_Language_Supervision.pdf
Learning_Versatile_Humanoid_Manipulation_with_Touch_Dreaming.pdf
LeJEPA_Provable_and_Scalable_Self-Supervised_Learning_Without_the_Heuristics.pdf
LeWorldModel_Stable_End-to-End_Joint-Embedding_Predictive_Architecture_from_Pixels.pdf
LLaMA_Open_and_Efficient_Foundation_Language_Models.pdf
LLM-JEPA_Large_Language_Models_Meet_Joint_Embedding_Predictive_Architectures.pdf
Mastering_Diverse_Domains_through_World_Models.pdf
Mathematical_Foundations_of_Reinforcement_Learning_赵世钰.pdf
MEM_Multi-Scale_Embodied_Memory_for_Vision_Language_Action_Models.pdf
Mimic_Intent_Not_Just_Trajectories.pdf
Motus2_A_Self-Evolving_General_World_Model_for_Dexterous_Manipulation.pdf
Neural_Machine_Translation_by_Jointly_Learning_to_Align_and_Translate.pdf
Octo_An_Open-Source_Generalist_Robot_Policy.pdf
Omega-0_A_Latent_Predictive_World_Action_Model_for_Concurrent_Humanoid_Loco-Manipulation.pdf
OmniVTLA_Vision_Tactile_Language_Action_Model.pdf
Online_RL_Fine-tuning_for_Flow-based_Vision-Language-Action_Models.pdf
Open_X-Embodiment_Robotic_Learning_Datasets_and_RT-X_Models.pdf
OpenHelix_A_Short_Survey_Empirical_Analysis_and_Open-Source_Dual-System_VLA_Model_for_Robotic_Manipulation.pdf
OpenVLA-OFT_Fine-Tuning_Vision-Language-Action_Models_Optimizing_Speed_and_Success.pdf
OpenVLA_An_Open-Source_Vision-Language-Action_Model.pdf
OpenWAM_An_Open_Modular_Exploration_Towards_Systematic_World-Action_Model_Pretraining.pdf
Orca_The_World_is_in_Your_Mind.pdf
PaliGemma_A_Versatile_3B_Vision-Language_Model_for_Transfer.pdf
pi0.5_A_VLA_Model_with_Open-World_Generalization.pdf
pi0.6_A_VLA_That_Learns_From_Experience.pdf
pi0.7_A_Steerable_Model_with_Emergent_Capabilities.pdf
pi0_A_Vision-Language-Action_Flow_Model.pdf
PP-Tac_Paper_Picking_Using_Tactile_Feedback.pdf
Precise_and_Dexterous_Robotic_Manipulation_via_Human-in-the-Loop_Reinforcement_Learning.pdf
Proximal_Policy_Optimization_Algorithms.pdf
Q-Learning.pdf
Qwen-RobotManip_Technical_Report_Alignment_Unlocks_Scale_for_Robotic_Manipulation_Foundation_Models.pdf
Qwen-RobotNav_Technical_Report_A_Scalable_Navigation_Model_Designed_for_an_Agentic_Navigation_System.pdf
Qwen-RobotWorld_Technical_Report_Unifying_Embodied_World_Modeling_through_Language-Conditioned_Video_Generation.pdf
Qwen-VLA_Unifying_Vision-Language-Action_Modeling_across_Tasks_Environments_and_Robot_Embodiments.pdf
R3M_A_Universal_Visual_Representation_for_Robot_Manipulation.pdf
RDT-1B_a_Diffusion_Foundation_Model_for_Bimanual_Manipulation.pdf
README.md
Real-Time_Execution_of_Action_Chunking_Flow_Policies.pdf
Retrieval-Augmented_Generation_for_Knowledge-Intensive_NLP_Tasks.pdf
ReWeight_Leveraging_Human_Data_for_VLA_Post-Training_via_Demonstration_Retrieval_and_Sample_Weighting.pdf
RL_Token_Bootstrapping_Online_RL_with_Vision-Language-Action_Models.pdf
RLinf_Flexible_and_Efficient_Large-scale_Reinforcement_Learning_via_Macro-to-Micro_Flow_Transformation.pdf
RoboMamba_Efficient_Vision-Language-Action_Model_for_Robotic_Reasoning_and_Manipulation.pdf
RT-1_Robotics_Transformer_for_Real-World_Control_at_Scale.pdf
RT-2_Vision-Language-Action_Models_Transfer_Web_Knowledge_to_Robotic_Control.pdf
SAM3D-Guided_Object-Centric_Representation_Alignment_for_Vision-Language-Action_Models.pdf
SARM_Stage-Aware_Reward_Modeling_for_Long_Horizon_Robot_Manipulation.pdf
Scaling_Laws_for_Neural_Language_Models.pdf
Self-supervised_perception_for_tactile_skin_covered_dexterous_hands.pdf
SigLIP_Sigmoid_Loss_for_Language_Image_Pre-Training.pdf
SmolVLA_A_Vision-Language-Action_Model_for_Affordable_and_Efficient_Robotics.pdf
Soft_Actor-Critic_Off-Policy_Maximum_Entropy_Deep_RL_with_a_Stochastic_Actor.pdf
SONIC_Supersizing_Motion_Tracking_for_Natural_Humanoid_Whole-Body_Control.pdf
Sparsh_Self-Supervised_Touch_Representations_for_Vision-Based_Tactile_Sensing.pdf
Spatial_Forcing_Implicit_Spatial_Representation_Alignment_for_Vision-language-action_Model.pdf
StarVLA_A_Lego_like_Codebase_for_Vision_Language_Action_Model_Developing.pdf
Switch_Transformers_Scaling_to_Trillion_Parameter_Models_with_Simple_and_Efficient_Sparsity.pdf
Tactile-VLA_Unlocking_VLA_Models_Physical_Knowledge_for_Tactile_Generalization.pdf
Tactile_Beyond_Pixels_Multisensory_Touch_Representations_for_Robot_Manipulation.pdf
TacVLA_Contact-Aware_Tactile_Fusion_for_Robust_VLA_Manipulation.pdf
TaF-VLA_Tactile-Force_Alignment_in_VLA_Models_for_Force-aware_Manipulation.pdf
tau0-VLA_a_Hierarchical_Robot_Foundation_Model_with_World-Model-Guided_Test-Time_Computation.pdf
Temporal_Difference_Learning_for_Model_Predictive_Control.pdf
TinyVLA_Towards_Fast_Data-Efficient_Vision-Language-Action_Models_for_Robotic_Manipulation.pdf
Towards_Human-Like_Manipulation_through_RL-Augmented_Teleoperation_and_Mixture-of-Dexterous-Experts_VLA.pdf
Training-Time_Action_Conditioning_for_Efficient_Real-Time_Chunking.pdf
Training_Compute-Optimal_Large_Language_Models.pdf
Training_Language_Models_to_Follow_Instructions_with_Human_Feedback.pdf
TurboVLA_Real-Time_Vision-Language-Action_Model_at_32_Hz_on_an_RTX_4090_with_Less_Than_1_GB_VRAM.pdf
TVL_A_Touch_Vision_and_Language_Dataset_for_Multimodal_Alignment.pdf
UMI_Universal_Manipulation_Interface_In-The-Wild_Robot_Teaching_Without_In-The-Wild_Robots.pdf
UniForce_A_Unified_Latent_Force_Model_for_Robot_Manipulation_with_Diverse_Tactile_Sensors.pdf
UniTouch_Binding_Touch_to_Everything_Learning_Unified_Multimodal_Tactile_Representations.pdf
Unleashing_Large-Scale_Video_Generative_Pre-training_for_Visual_Robot_Manipulation.pdf
V-JEPA_2_Self-Supervised_Video_Models_Enable_Understanding_Prediction_and_Planning.pdf
V-JEPA_Revisiting_Feature_Prediction_for_Learning_Visual_Representations_from_Video.pdf
Vision-Language-Action_Models_for_Bimanual_Manipulation_and_Real-World_Deployment_A_Comprehensive_Survey.pdf
Vision-Language-Action_Models_for_Robotics_A_Review_Towards_Real-World_Applications.pdf
Vision-Language_Foundation_Models_as_Effective_Robot_Imitators.pdf
Visual_Instruction_Tuning.pdf
VL-JEPA_Joint_Embedding_Predictive_Architecture_for_Vision-language.pdf
VLA-Adapter_An_Effective_Paradigm_for_Tiny-Scale_Vision-Language-Action_Model.pdf
VLA-JEPA_Enhancing_Vision-Language-Action_Model_with_Latent_World_Model.pdf
VLA-Touch_Enhancing_VLA_Models_with_Dual-Level_Tactile_Feedback.pdf
VLANeXt_Recipes_for_Building_Strong_VLA_Models.pdf
VTAM_Video-Tactile-Action_Models_for_Complex_Physical_Interaction_Beyond_VLAs.pdf
Wh0_Generative_World_Models_as_Scalable_Sources_of_Egocentric_Human_Hand_Manipulation_Data.pdf
World_Action_Models_The_Next_Frontier_in_Embodied_AI.pdf
World_Model_for_Robot_Learning_A_Comprehensive_Survey.pdf
X-VLA_Soft-Prompted_Transformer_as_Scalable_Cross-Embodiment_Vision-Language-Action_Model.pdf
Xiaomi-Robotics-1_Scaling_Vision-Language-Action_Models_with_over_100K_Hours_of_Real-World_Trajectories.pdf
Zero-WAM_In-Context_World-Action_Modeling_from_Human_Videos_for_Open-Ended_Task_Generalization.pdf