Views
No views yet
ce-amtic/ProcVLM-2B<progress>XX%</progress>1git clone https://github.com/ProcVLM/ProcVLM.git
2cd ProcVLM
3
4uv sync --python 3.10
5source .venv/bin/activate
6uv pip install flash-attn --no-build-isolation1python evqa/inference.py \
2 --model_path ce-amtic/ProcVLM-2B \
3 --video_path path/to/your/video.mp4 \
4 --output_path path/to/progress_predictions.jsonl \
5 --task "fold the red T-shirt" \
6 --window_size 8frame_index and its corresponding progress prediction.1python evqa/eval/visualize_progress_video.py \
2 --model_path ce-amtic/ProcVLM-2B \
3 --video_path path/to/your/video.mp4 \
4 --output_path path/to/progress_visualization.mp4 \
5 --task "fold the red T-shirt" \
6 --window_size 8infer_progress_from_video():1from evqa.inference import infer_progress_from_video
2
3records = infer_progress_from_video(
4 model_path="ce-amtic/ProcVLM-2B",
5 video_path="path/to/your/video.mp4",
6 task="fold the red T-shirt",
7 window_size=8,
8)
9
10for item in records:
11 print(item["frame_index"], item["progress"])frame_index: source video frame index;timestamp_sec: source video timestamp;window_frame_indices: frame indices used as the model input window;progress: parsed progress value in [0, 100];reasoning: model reasoning with the progress tag removed;model_output: raw model output.Given the recent observation and the task "{task}", first infer the remaining atomic actions required to complete the task. Then estimate the current completion percentage and output it as a float wrapped by <progress> tags.1To complete the task: Tower the blocks, the following steps are required:
21. Grasp the green block.
32. Place the green block onto the red block.
4Therefore, the estimated progress percentage is <progress>84.13%</progress>.1The task requires: Tower the blocks. Images show no block outside the tower, no further steps required.
2Therefore, the estimated progress percentage is <progress>100.00%</progress>.evqa.model.batch_chat_with_vllm():1from evqa.model import batch_chat_with_vllm
2
3outputs = batch_chat_with_vllm(
4 batch_items=[
5 {
6 "image": [
7 "frames/frame_000000.jpg",
8 "frames/frame_000010.jpg",
9 "frames/frame_000020.jpg",
10 ],
11 "conversations": [
12 {
13 "from": "human",
14 "value": 'Given the recent observation and the task "fold the red T-shirt", first infer the remaining atomic actions required to complete the task. Then estimate the current completion percentage and output it as a float wrapped by <progress> tags.',
15 }
16 ],
17 }
18 ],
19 model_path="ce-amtic/ProcVLM-2B",
20 max_new_tokens=1024,
21 temperature=0.0,
22 tp=1,
23)evqa/one-shot/lora_oneshot.sh;evqa/inference.py --use_lora.1@misc{feng2026procvlmlearningproceduregroundedprogress,
2 title={ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation},
3 author={Youhe Feng and Hansen Shi and Haoyang Li and Xinlei Guo and Yang Wang and Chengyang Zhang and Jinkai Zhang and Xiaohan Zhang and Jie Tang and Jing Zhang},
4 year={2026},
5 eprint={2605.08774},
6 archivePrefix={arXiv},
7 primaryClass={cs.RO},
8 url={https://arxiv.org/abs/2605.08774},
9}