Running experiment: plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025
 - kld_loss_weight=0.1
 - dice_loss_weight=8
 - teacher_ref_input=false
 - DICE_TYPE=focal_tversky
 - FOCAL_ALPHA=0.7
 - FOCAL_BETA=0.3
 - Using GPUs: 0,1
[2025-10-22 16:05:59,284] [INFO] [real_accelerator.py:110:get_accelerator] Setting ds_accelerator to cuda (auto detect)
[2025-10-22 16:06:03,178] [WARNING] [runner.py:196:fetch_hostfile] Unable to find hostfile, will proceed with training with local resources only.
[2025-10-22 16:06:03,178] [INFO] [runner.py:555:main] cmd = /shared/nas/data/m1/jk100/.conda/envs/llava/bin/python -u -m deepspeed.launcher.launch --world_info=eyJsb2NhbGhvc3QiOiBbMCwgMV19 --master_addr=127.0.0.1 --master_port=52367 --enable_each_rank_log=None plum_train_ds.py --version=liuhaotian/llava-llama-2-13b-chat-lightning-preview --dataset_dir=/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset --vision_pretrained=/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/weights/sam_vit_h_4b8939.pth --dataset=sem_seg||refer_seg||vqa||reason_seg --val_dataset=reason_seg|val --sample_rates=9,7,7,1 --batch_size=6 --grad_accumulation_steps 10 --use_bidir_bio --use_feedback_loop --ce_loss_weight=1.0 --dice_loss_weight=8 --bce_loss_weight=2.0 --kld_loss_weight=0.1 --seg_cls_loss_weight=2 --exp_name=plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025 --precision=bf16 --model_max_length 512 --dice_scale_factor 1000.0 --epochs 25 --train_mask_prompt_encoder --focal_tversky_alpha=0.7 --focal_tversky_beta=0.3 --bidir_dim_feedforward 2048 --auto_resume
[2025-10-22 16:06:04,322] [INFO] [real_accelerator.py:110:get_accelerator] Setting ds_accelerator to cuda (auto detect)
[2025-10-22 16:06:06,968] [INFO] [launch.py:145:main] WORLD INFO DICT: {'localhost': [0, 1]}
[2025-10-22 16:06:06,968] [INFO] [launch.py:151:main] nnodes=1, num_local_procs=2, node_rank=0
[2025-10-22 16:06:06,968] [INFO] [launch.py:162:main] global_rank_mapping=defaultdict(<class 'list'>, {'localhost': [0, 1]})
[2025-10-22 16:06:06,968] [INFO] [launch.py:163:main] dist_world_size=2
[2025-10-22 16:06:06,968] [INFO] [launch.py:165:main] Setting CUDA_VISIBLE_DEVICES=0,1
[2025-10-22 16:06:08,824] [INFO] [real_accelerator.py:110:get_accelerator] Setting ds_accelerator to cuda (auto detect)
[2025-10-22 16:06:08,824] [INFO] [real_accelerator.py:110:get_accelerator] Setting ds_accelerator to cuda (auto detect)
>> (train) run_name:  plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025_bidirbio_2048_maxlen512_epochs25_bsz6_lr0.0003_bidir_bio_feedback_loop_train_prompt_enc_srates_9_7_7_1
>> (train) run_ckpt_path:  plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025_bidirbio_2048_maxlen512_epochs25_bidir_bio_feedback_loop_train_prompt_enc_srates_9_7_7_1
>> (train) run_name:  plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025_bidirbio_2048_maxlen512_epochs25_bsz6_lr0.0003_bidir_bio_feedback_loop_train_prompt_enc_srates_9_7_7_1
>> (train) run_ckpt_path:  plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025_bidirbio_2048_maxlen512_epochs25_bidir_bio_feedback_loop_train_prompt_enc_srates_9_7_7_1
/home/jk100/.local/lib/python3.10/site-packages/huggingface_hub/file_download.py:945: FutureWarning: `resume_download` is deprecated and will be removed in version 1.0.0. Downloads always resume when possible. If you want to force a new download, use `force_download=True`.
  warnings.warn(
You are using the legacy behaviour of the <class 'transformers.models.llama.tokenization_llama.LlamaTokenizer'>. This means that tokens that come after special tokens will not be properly handled. We recommend you to read the related pull request available at https://github.com/huggingface/transformers/pull/24565
wandb: Currently logged in as: jeonghwankim123 (uiucnlp). Use `wandb login --relogin` to force relogin
wandb: wandb version 0.22.2 is available!  To upgrade, please run:
wandb:  $ pip install wandb --upgrade
wandb: Tracking run with wandb version 0.16.0
wandb: Run data is saved locally in /shared/nas2/jk100/partonomy_private/src/models/PLUM/wandb/run-20251022_160613-ltu7jgbz
wandb: Run `wandb offline` to turn off syncing.
wandb: Syncing run plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025_bidirbio_2048_maxlen512_epochs25_bsz6_lr0.0003_bidir_bio_feedback_loop_train_prompt_enc_srates_9_7_7_1
wandb: ⭐️ View project at https://wandb.ai/uiucnlp/plum-training
wandb: 🚀 View run at https://wandb.ai/uiucnlp/plum-training/runs/ltu7jgbz
/home/jk100/.local/lib/python3.10/site-packages/huggingface_hub/file_download.py:945: FutureWarning: `resume_download` is deprecated and will be removed in version 1.0.0. Downloads always resume when possible. If you want to force a new download, use `force_download=True`.
  warnings.warn(
You are using the legacy behaviour of the <class 'transformers.models.llama.tokenization_llama.LlamaTokenizer'>. This means that tokens that come after special tokens will not be properly handled. We recommend you to read the related pull request available at https://github.com/huggingface/transformers/pull/24565
/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/nn/modules/transformer.py:306: UserWarning: enable_nested_tensor is True, but self.use_nested_tensor is False because encoder_layer.self_attn.batch_first was not True(use batch_first for better inference performance)
  warnings.warn(f"enable_nested_tensor is True, but self.use_nested_tensor is False because {why_not_sparsity_fast_path}")
/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/nn/modules/transformer.py:306: UserWarning: enable_nested_tensor is True, but self.use_nested_tensor is False because encoder_layer.self_attn.batch_first was not True(use batch_first for better inference performance)
  warnings.warn(f"enable_nested_tensor is True, but self.use_nested_tensor is False because {why_not_sparsity_fast_path}")

Loading checkpoint shards:   0%|          | 0/3 [00:00<?, ?it/s]
Loading checkpoint shards:  33%|███▎      | 1/3 [00:06<00:12,  6.24s/it]
Loading checkpoint shards:   0%|          | 0/3 [00:00<?, ?it/s]
Loading checkpoint shards:  67%|██████▋   | 2/3 [00:12<00:06,  6.27s/it]
Loading checkpoint shards:  33%|███▎      | 1/3 [00:06<00:12,  6.13s/it]
Loading checkpoint shards: 100%|██████████| 3/3 [00:16<00:00,  5.32s/it]
Loading checkpoint shards: 100%|██████████| 3/3 [00:16<00:00,  5.58s/it]
Some weights of PLUMForCausalLM were not initialized from the model checkpoint at liuhaotian/llava-llama-2-13b-chat-lightning-preview and are newly initialized: ['model.visual_model.mask_decoder.transformer.layers.1.norm1.weight', 'model.visual_model.prompt_encoder.not_a_point_embed.weight', 'mask_pooler.final_conv.weight', 'model.visual_model.image_encoder.blocks.25.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.28.attn.proj.bias', 'model.visual_model.image_encoder.blocks.1.norm2.weight', 'model.visual_model.image_encoder.blocks.14.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.31.norm2.bias', 'model.visual_model.image_encoder.blocks.27.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.12.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.weight', 'model.visual_model.image_encoder.blocks.6.norm1.weight', 'mask_pooler.block1.0.1.weight', 'model.visual_model.prompt_encoder.point_embeddings.2.weight', 'model.visual_model.image_encoder.blocks.14.norm1.bias', 'model.visual_model.image_encoder.blocks.23.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.weight', 'model.visual_model.image_encoder.blocks.10.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.7.attn.proj.bias', 'model.visual_model.image_encoder.blocks.7.norm2.weight', 'model.visual_model.image_encoder.blocks.13.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.15.attn.proj.weight', 'model.visual_model.image_encoder.blocks.4.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.13.attn.qkv.bias', 'model.visual_model.prompt_encoder.mask_downscaling.0.weight', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.weight', 'model.visual_model.image_encoder.blocks.30.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.12.attn.qkv.bias', 'mask_pooler.final_conv.bias', 'model.visual_model.image_encoder.blocks.25.norm1.bias', 'model.visual_model.mask_decoder.iou_prediction_head.layers.2.bias', 'model.visual_model.prompt_encoder.point_embeddings.1.weight', 'model.visual_model.image_encoder.blocks.20.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.9.attn.proj.bias', 'model.visual_model.image_encoder.blocks.24.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.17.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.26.attn.qkv.weight', 'model.visual_model.image_encoder.neck.3.bias', 'model.visual_model.image_encoder.blocks.8.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.10.attn.proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.norm4.bias', 'model.visual_model.image_encoder.blocks.17.norm2.bias', 'model.visual_model.image_encoder.neck.1.bias', 'model.visual_model.image_encoder.blocks.29.attn.qkv.weight', 'model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.2.attn.proj.weight', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.bias', 'model.visual_model.image_encoder.blocks.25.attn.proj.bias', 'model.visual_model.image_encoder.blocks.22.norm2.bias', 'model.visual_model.image_encoder.blocks.3.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.norm_final_attn.bias', 'model.visual_model.image_encoder.blocks.31.attn.rel_pos_w', 'model.visual_model.prompt_encoder.mask_downscaling.6.bias', 'model.visual_model.image_encoder.blocks.3.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.17.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.8.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.17.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.0.norm1.bias', 'model.visual_model.image_encoder.blocks.29.norm2.bias', 'model.visual_model.image_encoder.blocks.30.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.18.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.bias', 'model.visual_model.image_encoder.blocks.0.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.7.mlp.lin2.bias', 'mask_pooler.prompt_encoder.not_a_point_embed.weight', 'mask_pooler.prompt_encoder.mask_downscaling.1.bias', 'model.visual_model.image_encoder.blocks.2.norm1.bias', 'model.visual_model.image_encoder.blocks.7.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.24.norm1.weight', 'model.visual_model.image_encoder.blocks.20.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.bias', 'model.visual_model.image_encoder.blocks.29.norm1.weight', 'model.visual_model.image_encoder.blocks.20.norm2.weight', 'model.visual_model.image_encoder.blocks.29.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.11.attn.rel_pos_w', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.weight', 'model.visual_model.image_encoder.blocks.21.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.8.mlp.lin2.bias', 'mask_pooler.attn.out_proj.weight', 'model.visual_model.image_encoder.blocks.3.norm1.bias', 'model.visual_model.image_encoder.blocks.14.norm2.weight', 'model.visual_model.image_encoder.blocks.19.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.12.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.17.norm1.bias', 'model.visual_model.image_encoder.blocks.17.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.30.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.7.norm1.bias', 'model.visual_model.image_encoder.blocks.10.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.11.norm2.bias', 'model.visual_model.image_encoder.blocks.16.attn.proj.bias', 'model.visual_model.image_encoder.blocks.3.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.17.norm1.weight', 'model.visual_model.image_encoder.blocks.13.mlp.lin2.bias', 'model.visual_model.image_encoder.neck.0.weight', 'model.visual_model.image_encoder.blocks.9.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.13.attn.rel_pos_w', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.bias', 'model.visual_model.mask_decoder.transformer.layers.1.norm1.bias', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.weight', 'model.visual_model.image_encoder.blocks.21.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.0.attn.proj.weight', 'model.visual_model.image_encoder.blocks.8.attn.rel_pos_w', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.weight', 'model.visual_model.image_encoder.blocks.14.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.8.attn.proj.weight', 'model.visual_model.image_encoder.blocks.1.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.9.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.22.mlp.lin2.weight', 'model.visual_model.prompt_encoder.point_embeddings.3.weight', 'model.visual_model.image_encoder.blocks.5.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.15.norm2.weight', 'model.visual_model.image_encoder.blocks.22.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.26.attn.proj.weight', 'model.visual_model.image_encoder.blocks.31.mlp.lin2.bias', 'mask_pooler.prompt_encoder.mask_downscaling.4.weight', 'model.visual_model.image_encoder.blocks.2.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.6.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.5.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.23.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.6.norm2.bias', 'model.visual_model.image_encoder.blocks.22.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.13.attn.proj.weight', 'model.visual_model.image_encoder.blocks.0.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.1.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.9.norm2.weight', 'model.visual_model.image_encoder.blocks.26.attn.rel_pos_w', 'model.token_to_mask_fcs.0.2.bias', 'model.visual_model.image_encoder.blocks.5.norm2.weight', 'model.token_to_mask_fcs.0.2.weight', 'model.visual_model.image_encoder.blocks.19.norm1.weight', 'model.visual_model.image_encoder.blocks.5.norm1.weight', 'model.visual_model.image_encoder.blocks.23.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.12.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.16.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.9.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.15.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.4.attn.proj.weight', 'model.visual_model.image_encoder.blocks.2.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.25.norm1.weight', 'model.visual_model.image_encoder.blocks.7.attn.rel_pos_h', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.weight', 'model.visual_model.image_encoder.blocks.31.attn.proj.bias', 'model.visual_model.image_encoder.blocks.21.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.29.attn.proj.weight', 'model.visual_model.image_encoder.blocks.18.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.15.norm1.weight', 'model.visual_model.prompt_encoder.mask_downscaling.0.bias', 'model.visual_model.image_encoder.blocks.11.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.31.attn.qkv.bias', 'mask_pooler.attn.in_proj_bias', 'model.visual_model.image_encoder.blocks.10.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.29.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.22.norm1.bias', 'model.visual_model.image_encoder.blocks.10.norm1.bias', 'model.visual_model.image_encoder.blocks.28.norm1.weight', 'model.visual_model.image_encoder.blocks.0.norm2.weight', 'model.visual_model.image_encoder.blocks.27.norm2.weight', 'model.visual_model.image_encoder.blocks.14.norm2.bias', 'model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.weight', 'model.visual_model.image_encoder.blocks.2.norm2.weight', 'model.visual_model.image_encoder.blocks.15.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.4.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.2.attn.proj.bias', 'model.visual_model.image_encoder.blocks.7.norm1.weight', 'model.visual_model.image_encoder.blocks.25.mlp.lin2.bias', 'model.visual_model.mask_decoder.iou_prediction_head.layers.1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.weight', 'mask_pooler.mask_modulator.projections.0.weight', 'mask_pooler.attn.out_proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.bias', 'model.visual_model.image_encoder.blocks.31.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.weight', 'model.visual_model.image_encoder.blocks.28.attn.rel_pos_h', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.bias', 'model.visual_model.image_encoder.blocks.25.norm2.bias', 'mask_pooler.prompt_encoder.mask_downscaling.3.weight', 'model.visual_model.image_encoder.blocks.1.attn.qkv.bias', 'mask_pooler.prompt_encoder.mask_downscaling.4.bias', 'model.visual_model.image_encoder.blocks.15.attn.qkv.bias', 'model.visual_model.mask_decoder.iou_prediction_head.layers.0.weight', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.bias', 'model.visual_model.image_encoder.blocks.14.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.23.attn.qkv.weight', 'mask_pooler.prompt_encoder.mask_downscaling.6.bias', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.bias', 'model.visual_model.image_encoder.blocks.9.norm1.bias', 'model.visual_model.image_encoder.blocks.25.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.12.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.24.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.11.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.24.norm2.bias', 'model.visual_model.prompt_encoder.mask_downscaling.4.weight', 'model.visual_model.image_encoder.blocks.21.attn.proj.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.bias', 'model.visual_model.image_encoder.blocks.17.attn.proj.weight', 'model.visual_model.image_encoder.blocks.11.attn.rel_pos_h', 'model.visual_model.prompt_encoder.pe_layer.positional_encoding_gaussian_matrix', 'model.visual_model.image_encoder.blocks.18.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.bias', 'model.visual_model.image_encoder.blocks.27.mlp.lin2.bias', 'bio_encoder.encoder.layers.0.norm2.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.norm3.bias', 'model.visual_model.image_encoder.blocks.28.attn.qkv.weight', 'model.visual_model.image_encoder.neck.2.weight', 'model.visual_model.image_encoder.blocks.4.attn.proj.bias', 'model.visual_model.image_encoder.blocks.17.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.5.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.19.mlp.lin1.bias', 'mask_pooler.prompt_encoder.mask_downscaling.1.weight', 'model.visual_model.image_encoder.blocks.23.norm1.bias', 'model.visual_model.image_encoder.blocks.21.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.weight', 'model.visual_model.image_encoder.blocks.20.norm2.bias', 'model.visual_model.image_encoder.blocks.2.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.5.attn.proj.weight', 'model.visual_model.image_encoder.blocks.12.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.20.mlp.lin2.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.bias', 'model.visual_model.image_encoder.blocks.12.attn.proj.weight', 'model.visual_model.image_encoder.blocks.25.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.17.norm2.weight', 'model.visual_model.image_encoder.blocks.10.attn.qkv.weight', 'mask_pooler.prompt_encoder.mask_downscaling.0.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.weight', 'model.visual_model.image_encoder.blocks.31.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.26.norm1.bias', 'model.visual_model.image_encoder.blocks.31.attn.rel_pos_h', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.weight', 'model.visual_model.image_encoder.blocks.24.mlp.lin1.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.weight', 'model.visual_model.image_encoder.blocks.4.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.weight', 'model.visual_model.image_encoder.blocks.22.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.15.attn.proj.bias', 'model.visual_model.image_encoder.neck.1.weight', 'model.visual_model.image_encoder.blocks.30.attn.proj.weight', 'model.token_to_mask_fcs.0.0.bias', 'bio_encoder.encoder.layers.0.self_attn.out_proj.weight', 'model.visual_model.image_encoder.blocks.18.attn.rel_pos_w', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.weight', 'model.visual_model.image_encoder.blocks.27.norm1.bias', 'model.visual_model.image_encoder.blocks.13.norm2.weight', 'model.visual_model.image_encoder.blocks.30.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.26.mlp.lin2.weight', 'model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.bias', 'model.visual_model.prompt_encoder.no_mask_embed.weight', 'model.visual_model.image_encoder.blocks.9.norm1.weight', 'model.visual_model.image_encoder.blocks.27.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.30.norm2.bias', 'model.visual_model.image_encoder.blocks.1.attn.proj.bias', 'model.visual_model.image_encoder.blocks.1.norm2.bias', 'bio_encoder.encoder.layers.0.linear1.bias', 'model.visual_model.image_encoder.blocks.3.attn.qkv.bias', 'mask_pooler.prompt_encoder.mask_downscaling.6.weight', 'mask_pooler.mask_modulator.projections.2.bias', 'model.visual_model.image_encoder.blocks.26.mlp.lin1.bias', 'mask_pooler.block2.0.3.weight', 'model.visual_model.image_encoder.blocks.19.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.29.attn.qkv.bias', 'model.visual_model.mask_decoder.iou_prediction_head.layers.2.weight', 'model.visual_model.image_encoder.blocks.16.norm1.bias', 'model.visual_model.image_encoder.blocks.17.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.14.norm1.weight', 'model.visual_model.image_encoder.blocks.29.mlp.lin2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.weight', 'model.visual_model.image_encoder.blocks.18.attn.qkv.weight', 'mask_pooler.mask_modulator.projections.0.bias', 'model.visual_model.image_encoder.blocks.7.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.8.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.19.norm1.bias', 'model.visual_model.image_encoder.blocks.16.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.0.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.24.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.1.attn.rel_pos_h', 'bio_encoder.encoder.layers.0.self_attn.out_proj.bias', 'model.visual_model.image_encoder.blocks.26.norm2.bias', 'model.visual_model.image_encoder.pos_embed', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.bias', 'model.visual_model.image_encoder.blocks.3.attn.proj.weight', 'model.visual_model.image_encoder.blocks.22.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.7.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.29.norm1.bias', 'model.visual_model.image_encoder.blocks.11.norm1.weight', 'model.visual_model.image_encoder.blocks.0.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.15.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.bias', 'model.visual_model.image_encoder.blocks.18.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.24.attn.qkv.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.bias', 'model.visual_model.image_encoder.blocks.3.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.6.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.4.norm2.bias', 'model.visual_model.image_encoder.blocks.23.attn.proj.bias', 'model.visual_model.image_encoder.blocks.28.mlp.lin2.weight', 'bio_encoder.encoder.layers.0.norm1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.weight', 'model.visual_model.image_encoder.blocks.23.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.10.norm2.bias', 'model.visual_model.image_encoder.blocks.28.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.8.attn.qkv.weight', 'model.visual_model.mask_decoder.transformer.layers.1.norm2.weight', 'model.visual_model.image_encoder.blocks.20.norm1.weight', 'model.visual_model.image_encoder.blocks.29.attn.proj.bias', 'mask_pooler.mask_modulator.projections.2.weight', 'model.visual_model.image_encoder.blocks.1.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.18.norm2.weight', 'mask_pooler.prompt_encoder.no_mask_embed.weight', 'model.visual_model.image_encoder.blocks.4.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.25.norm2.weight', 'mask_pooler.block1.0.0.weight', 'model.visual_model.image_encoder.blocks.1.attn.qkv.weight', 'model.text_hidden_fcs.0.2.weight', 'bio_encoder.encoder.layers.0.norm1.bias', 'model.visual_model.image_encoder.blocks.16.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.18.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.16.norm2.bias', 'model.visual_model.image_encoder.blocks.7.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.weight', 'model.visual_model.image_encoder.blocks.20.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.10.norm2.weight', 'model.visual_model.image_encoder.blocks.24.attn.proj.bias', 'model.visual_model.image_encoder.blocks.14.attn.proj.bias', 'model.visual_model.image_encoder.blocks.19.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.8.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.21.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.22.attn.proj.bias', 'model.visual_model.mask_decoder.output_upscaling.3.weight', 'model.visual_model.image_encoder.blocks.5.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.1.norm1.weight', 'model.visual_model.image_encoder.blocks.12.norm1.weight', 'model.visual_model.image_encoder.blocks.30.attn.rel_pos_h', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.bias', 'model.visual_model.image_encoder.blocks.30.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.bias', 'model.token_to_mask_fcs.0.0.weight', 'model.visual_model.image_encoder.blocks.22.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.bias', 'model.visual_model.image_encoder.blocks.30.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.29.norm2.weight', 'mask_pooler.prompt_encoder.point_embeddings.0.weight', 'model.visual_model.image_encoder.blocks.13.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.13.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.2.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.13.norm2.bias', 'model.visual_model.image_encoder.blocks.19.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.bias', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.weight', 'bio_encoder.encoder.layers.0.self_attn.in_proj_weight', 'model.visual_model.image_encoder.blocks.2.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.16.mlp.lin1.weight', 'model.visual_model.mask_decoder.mask_tokens.weight', 'model.visual_model.image_encoder.blocks.3.norm2.bias', 'model.visual_model.image_encoder.blocks.3.norm1.weight', 'model.visual_model.image_encoder.blocks.6.norm2.weight', 'model.visual_model.image_encoder.blocks.14.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.9.attn.qkv.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.bias', 'model.visual_model.image_encoder.blocks.15.norm1.bias', 'model.visual_model.image_encoder.blocks.5.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.30.norm2.weight', 'model.visual_model.image_encoder.blocks.26.mlp.lin1.weight', 'mask_pooler.attn.in_proj_weight', 'model.visual_model.image_encoder.blocks.27.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.8.norm1.weight', 'model.visual_model.image_encoder.blocks.22.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.31.norm1.weight', 'bio_encoder.encoder.layers.0.self_attn.in_proj_bias', 'model.visual_model.mask_decoder.transformer.layers.1.norm3.weight', 'model.visual_model.image_encoder.blocks.4.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.14.attn.rel_pos_w', 'mask_pooler.prompt_encoder.mask_downscaling.0.weight', 'model.visual_model.image_encoder.blocks.16.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.23.mlp.lin2.weight', 'mask_pooler.prompt_encoder.point_embeddings.3.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.bias', 'mask_pooler.block2.0.4.weight', 'model.visual_model.image_encoder.blocks.3.attn.rel_pos_h', 'model.visual_model.mask_decoder.transformer.layers.0.norm1.bias', 'model.visual_model.image_encoder.blocks.8.norm1.bias', 'model.visual_model.image_encoder.blocks.4.norm1.weight', 'model.visual_model.image_encoder.blocks.31.attn.proj.weight', 'model.visual_model.image_encoder.blocks.5.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.21.norm1.weight', 'model.visual_model.image_encoder.blocks.19.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.23.attn.proj.weight', 'model.visual_model.image_encoder.blocks.28.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.18.attn.proj.weight', 'model.visual_model.image_encoder.blocks.24.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.1.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.22.attn.proj.weight', 'bio_encoder.encoder.layers.0.linear2.bias', 'model.visual_model.mask_decoder.transformer.layers.0.norm4.bias', 'model.visual_model.image_encoder.blocks.13.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.11.attn.proj.bias', 'model.visual_model.image_encoder.blocks.6.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.0.norm1.weight', 'model.visual_model.image_encoder.blocks.5.norm1.bias', 'model.visual_model.image_encoder.blocks.21.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.bias', 'model.text_hidden_fcs.0.0.bias', 'model.visual_model.image_encoder.blocks.21.attn.proj.weight', 'model.visual_model.image_encoder.blocks.28.attn.proj.weight', 'model.visual_model.prompt_encoder.mask_downscaling.1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.bias', 'model.visual_model.image_encoder.blocks.24.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.25.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.25.attn.proj.weight', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.weight', 'model.visual_model.image_encoder.blocks.15.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.2.norm2.bias', 'model.visual_model.image_encoder.blocks.6.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.12.norm1.bias', 'model.visual_model.image_encoder.blocks.14.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.28.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.5.norm2.bias', 'model.visual_model.image_encoder.blocks.0.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.23.norm2.weight', 'model.visual_model.image_encoder.blocks.12.norm2.bias', 'model.visual_model.image_encoder.blocks.14.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.8.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.19.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.30.norm1.bias', 'model.visual_model.image_encoder.blocks.18.attn.proj.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.weight', 'model.visual_model.image_encoder.blocks.8.attn.proj.bias', 'model.visual_model.image_encoder.blocks.15.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.7.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.4.norm2.weight', 'model.visual_model.mask_decoder.iou_token.weight', 'model.visual_model.image_encoder.blocks.4.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.10.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm3.weight', 'model.visual_model.image_encoder.blocks.1.norm1.bias', 'model.visual_model.image_encoder.blocks.27.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.28.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.9.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.bias', 'model.visual_model.image_encoder.blocks.0.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.10.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.31.norm1.bias', 'model.visual_model.image_encoder.blocks.0.attn.proj.bias', 'model.visual_model.mask_decoder.iou_prediction_head.layers.1.weight', 'mask_pooler.block1.0.0.bias', 'model.visual_model.image_encoder.blocks.8.norm2.bias', 'model.visual_model.image_encoder.blocks.30.attn.proj.bias', 'model.visual_model.image_encoder.blocks.19.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm3.bias', 'model.visual_model.image_encoder.blocks.2.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.9.attn.proj.weight', 'model.visual_model.image_encoder.blocks.24.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.norm2.bias', 'model.visual_model.image_encoder.blocks.12.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.9.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.16.norm1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.bias', 'model.visual_model.image_encoder.blocks.18.norm1.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.weight', 'model.visual_model.image_encoder.blocks.21.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.1.attn.proj.weight', 'model.visual_model.image_encoder.blocks.20.attn.rel_pos_h', 'model.visual_model.mask_decoder.transformer.layers.1.norm4.weight', 'mask_pooler.prompt_encoder.point_embeddings.1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.bias', 'model.visual_model.mask_decoder.output_upscaling.0.bias', 'model.visual_model.image_encoder.blocks.2.norm1.weight', 'model.visual_model.image_encoder.patch_embed.proj.weight', 'model.visual_model.image_encoder.blocks.10.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.31.mlp.lin1.weight', 'model.visual_model.prompt_encoder.mask_downscaling.3.bias', 'model.visual_model.image_encoder.blocks.6.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.13.mlp.lin2.weight', 'model.visual_model.image_encoder.neck.3.weight', 'model.visual_model.image_encoder.blocks.2.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.26.norm1.weight', 'model.visual_model.image_encoder.blocks.9.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.20.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.bias', 'model.visual_model.image_encoder.blocks.12.attn.rel_pos_h', 'model.visual_model.mask_decoder.output_upscaling.1.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.weight', 'model.visual_model.image_encoder.blocks.2.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.14.attn.proj.weight', 'model.visual_model.image_encoder.blocks.23.norm2.bias', 'model.visual_model.image_encoder.blocks.24.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.norm2.weight', 'mask_pooler.prompt_encoder.mask_downscaling.3.bias', 'model.visual_model.image_encoder.blocks.4.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.17.attn.proj.bias', 'model.visual_model.image_encoder.blocks.27.attn.proj.weight', 'model.visual_model.image_encoder.blocks.21.attn.qkv.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.weight', 'model.visual_model.image_encoder.blocks.26.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.9.norm2.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.bias', 'model.visual_model.image_encoder.blocks.27.attn.proj.bias', 'model.visual_model.image_encoder.blocks.28.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.27.attn.rel_pos_h', 'mask_pooler.prompt_encoder.pe_layer.positional_encoding_gaussian_matrix', 'model.visual_model.image_encoder.blocks.10.attn.proj.weight', 'model.visual_model.image_encoder.blocks.19.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.15.norm2.bias', 'model.visual_model.image_encoder.blocks.20.attn.proj.bias', 'mask_pooler.mask_modulator.projections.1.bias', 'model.visual_model.image_encoder.blocks.22.mlp.lin1.bias', 'model.text_hidden_fcs.0.0.weight', 'model.visual_model.image_encoder.blocks.3.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.25.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.30.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.22.attn.qkv.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.weight', 'model.visual_model.image_encoder.patch_embed.proj.bias', 'model.visual_model.image_encoder.blocks.25.attn.rel_pos_h', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.bias', 'mask_pooler.block2.0.4.bias', 'model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.weight', 'mask_pooler.prompt_encoder.point_embeddings.2.weight', 'model.visual_model.mask_decoder.output_upscaling.3.bias', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.weight', 'model.visual_model.image_encoder.blocks.4.attn.qkv.bias', 'model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.weight', 'model.visual_model.image_encoder.blocks.10.norm1.weight', 'model.visual_model.mask_decoder.transformer.norm_final_attn.weight', 'model.visual_model.image_encoder.blocks.28.norm2.bias', 'model.visual_model.image_encoder.blocks.26.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.28.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm1.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.weight', 'bio_encoder.encoder.layers.0.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm4.weight', 'model.visual_model.image_encoder.blocks.15.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.11.norm2.weight', 'model.visual_model.image_encoder.blocks.18.norm1.weight', 'model.visual_model.image_encoder.blocks.29.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.0.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.23.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.13.attn.proj.bias', 'model.visual_model.image_encoder.blocks.6.attn.proj.weight', 'model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.17.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.19.attn.proj.weight', 'model.visual_model.image_encoder.blocks.24.norm2.weight', 'model.visual_model.prompt_encoder.point_embeddings.0.weight', 'model.visual_model.image_encoder.blocks.14.attn.qkv.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.weight', 'model.visual_model.image_encoder.blocks.30.norm1.weight', 'model.visual_model.image_encoder.blocks.15.mlp.lin1.weight', 'model.visual_model.prompt_encoder.mask_downscaling.6.weight', 'model.visual_model.image_encoder.blocks.22.norm1.weight', 'model.visual_model.image_encoder.blocks.26.attn.proj.bias', 'model.visual_model.image_encoder.blocks.9.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.19.norm2.bias', 'mask_pooler.block2.0.3.bias', 'model.visual_model.prompt_encoder.mask_downscaling.3.weight', 'model.visual_model.image_encoder.blocks.4.attn.rel_pos_w', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.weight', 'model.visual_model.image_encoder.blocks.5.attn.proj.bias', 'model.visual_model.image_encoder.blocks.3.norm2.weight', 'model.visual_model.image_encoder.blocks.6.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.3.attn.proj.bias', 'model.visual_model.image_encoder.blocks.5.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.29.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.16.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.20.norm1.bias', 'mask_pooler.mask_modulator.projections.1.weight', 'model.visual_model.image_encoder.blocks.25.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.27.norm2.bias', 'bio_encoder.encoder.layers.0.linear1.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.weight', 'model.visual_model.image_encoder.blocks.27.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.8.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm2.bias', 'mask_pooler.queries', 'model.visual_model.image_encoder.blocks.16.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.29.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.18.norm2.bias', 'model.visual_model.image_encoder.blocks.3.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.27.mlp.lin2.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.weight', 'model.visual_model.image_encoder.blocks.20.attn.proj.weight', 'model.visual_model.image_encoder.blocks.7.norm2.bias', 'model.visual_model.image_encoder.blocks.10.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.16.attn.proj.weight', 'model.visual_model.image_encoder.blocks.0.norm2.bias', 'model.visual_model.image_encoder.blocks.21.norm2.bias', 'model.visual_model.image_encoder.blocks.7.attn.proj.weight', 'model.visual_model.image_encoder.blocks.31.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.31.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.weight', 'model.visual_model.image_encoder.blocks.18.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.13.norm1.bias', 'model.visual_model.image_encoder.blocks.26.norm2.weight', 'model.visual_model.image_encoder.blocks.11.attn.proj.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.bias', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.bias', 'model.visual_model.image_encoder.blocks.11.mlp.lin2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.bias', 'model.text_hidden_fcs.0.2.bias', 'model.visual_model.image_encoder.blocks.11.attn.qkv.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.weight', 'model.visual_model.image_encoder.blocks.17.attn.rel_pos_w', 'model.visual_model.mask_decoder.output_upscaling.0.weight', 'model.visual_model.image_encoder.blocks.6.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.20.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.6.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.12.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.21.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.bias', 'model.visual_model.image_encoder.blocks.27.norm1.weight', 'model.visual_model.image_encoder.blocks.5.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.13.norm1.weight', 'model.visual_model.mask_decoder.output_upscaling.1.bias', 'model.visual_model.image_encoder.blocks.16.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.11.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.6.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.20.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.28.norm1.bias', 'mask_pooler.block1.0.1.bias', 'model.visual_model.mask_decoder.iou_prediction_head.layers.0.bias', 'bio_encoder.encoder.layers.0.linear2.weight', 'model.visual_model.image_encoder.blocks.0.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.23.norm1.weight', 'model.visual_model.image_encoder.blocks.26.attn.rel_pos_h', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.weight', 'model.visual_model.image_encoder.blocks.23.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.1.mlp.lin2.weight', 'model.visual_model.prompt_encoder.mask_downscaling.1.bias', 'model.visual_model.image_encoder.blocks.7.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.11.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.16.norm2.weight', 'model.visual_model.image_encoder.blocks.24.attn.proj.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.bias', 'model.visual_model.image_encoder.blocks.19.attn.proj.bias', 'model.visual_model.prompt_encoder.mask_downscaling.4.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.bias', 'model.visual_model.image_encoder.blocks.12.attn.proj.bias', 'model.visual_model.image_encoder.blocks.6.attn.proj.bias', 'model.visual_model.image_encoder.blocks.21.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.11.norm1.bias']
You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.

Loading checkpoint shards:  67%|██████▋   | 2/3 [00:12<00:06,  6.03s/it]
Loading checkpoint shards: 100%|██████████| 3/3 [00:16<00:00,  5.14s/it]
Loading checkpoint shards: 100%|██████████| 3/3 [00:16<00:00,  5.39s/it]
Some weights of PLUMForCausalLM were not initialized from the model checkpoint at liuhaotian/llava-llama-2-13b-chat-lightning-preview and are newly initialized: ['model.visual_model.image_encoder.blocks.15.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.12.norm1.bias', 'model.visual_model.image_encoder.blocks.25.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.15.norm1.bias', 'mask_pooler.final_conv.weight', 'mask_pooler.prompt_encoder.point_embeddings.1.weight', 'model.visual_model.image_encoder.blocks.30.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.10.norm1.weight', 'model.visual_model.image_encoder.blocks.26.norm2.weight', 'model.visual_model.image_encoder.blocks.0.norm1.weight', 'model.visual_model.image_encoder.blocks.3.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.18.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.23.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.bias', 'model.visual_model.image_encoder.blocks.9.attn.proj.bias', 'model.visual_model.image_encoder.blocks.18.norm1.bias', 'model.visual_model.image_encoder.blocks.27.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.17.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.13.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.15.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.16.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.27.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.31.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.7.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.24.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.20.norm2.bias', 'model.visual_model.image_encoder.blocks.21.norm2.weight', 'model.visual_model.image_encoder.blocks.27.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.31.attn.qkv.weight', 'mask_pooler.block2.0.3.weight', 'model.visual_model.image_encoder.blocks.27.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.30.attn.qkv.bias', 'model.visual_model.image_encoder.patch_embed.proj.bias', 'model.visual_model.image_encoder.blocks.17.attn.qkv.weight', 'model.visual_model.image_encoder.neck.3.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.bias', 'model.visual_model.image_encoder.blocks.11.norm1.weight', 'model.visual_model.image_encoder.blocks.19.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.7.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.norm3.weight', 'model.visual_model.image_encoder.blocks.11.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.28.norm1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.bias', 'model.text_hidden_fcs.0.0.weight', 'model.visual_model.image_encoder.blocks.24.attn.qkv.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.bias', 'model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.0.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.19.attn.proj.bias', 'model.visual_model.image_encoder.blocks.1.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.4.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.1.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.2.norm1.bias', 'mask_pooler.block2.0.3.bias', 'model.visual_model.image_encoder.blocks.1.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.1.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.25.attn.proj.weight', 'model.visual_model.image_encoder.blocks.10.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.16.mlp.lin2.weight', 'model.visual_model.prompt_encoder.mask_downscaling.0.bias', 'mask_pooler.prompt_encoder.mask_downscaling.4.bias', 'model.visual_model.image_encoder.blocks.14.norm1.weight', 'model.visual_model.image_encoder.blocks.9.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.16.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.20.norm2.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.weight', 'model.visual_model.mask_decoder.iou_prediction_head.layers.2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.bias', 'model.visual_model.image_encoder.neck.1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm4.weight', 'model.visual_model.image_encoder.blocks.22.attn.proj.bias', 'model.visual_model.image_encoder.blocks.5.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.17.norm1.bias', 'model.visual_model.image_encoder.blocks.6.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.8.attn.qkv.weight', 'mask_pooler.prompt_encoder.mask_downscaling.1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.weight', 'model.visual_model.image_encoder.blocks.8.mlp.lin2.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.bias', 'model.visual_model.image_encoder.blocks.1.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.bias', 'mask_pooler.prompt_encoder.no_mask_embed.weight', 'model.visual_model.image_encoder.blocks.18.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.17.norm2.bias', 'model.visual_model.image_encoder.blocks.14.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.0.norm2.weight', 'model.visual_model.image_encoder.blocks.17.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.26.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.12.norm2.weight', 'bio_encoder.encoder.layers.0.linear1.bias', 'model.visual_model.image_encoder.blocks.26.attn.qkv.weight', 'mask_pooler.prompt_encoder.point_embeddings.3.weight', 'model.visual_model.image_encoder.blocks.5.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.31.attn.proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.bias', 'mask_pooler.prompt_encoder.pe_layer.positional_encoding_gaussian_matrix', 'model.visual_model.image_encoder.blocks.1.attn.proj.weight', 'model.visual_model.image_encoder.blocks.3.norm1.weight', 'model.visual_model.image_encoder.neck.3.bias', 'model.visual_model.image_encoder.blocks.8.attn.qkv.bias', 'model.visual_model.prompt_encoder.pe_layer.positional_encoding_gaussian_matrix', 'model.visual_model.mask_decoder.output_upscaling.1.bias', 'model.visual_model.image_encoder.blocks.0.norm1.bias', 'mask_pooler.mask_modulator.projections.2.weight', 'model.visual_model.image_encoder.blocks.12.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.22.norm1.weight', 'model.visual_model.image_encoder.blocks.22.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.4.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.24.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.10.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.4.attn.proj.bias', 'model.visual_model.image_encoder.blocks.3.norm2.bias', 'model.visual_model.mask_decoder.output_upscaling.3.bias', 'model.visual_model.image_encoder.blocks.31.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.18.attn.qkv.bias', 'model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.2.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.12.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.2.attn.rel_pos_w', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.weight', 'model.visual_model.image_encoder.blocks.4.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.30.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.18.norm2.weight', 'model.visual_model.image_encoder.neck.0.weight', 'model.visual_model.image_encoder.blocks.21.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.21.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.20.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.13.norm1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.bias', 'model.visual_model.image_encoder.blocks.20.attn.proj.bias', 'model.visual_model.image_encoder.pos_embed', 'model.visual_model.image_encoder.blocks.5.attn.proj.weight', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.weight', 'model.visual_model.image_encoder.blocks.0.norm2.bias', 'model.visual_model.image_encoder.blocks.22.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.31.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.17.norm1.weight', 'model.visual_model.image_encoder.blocks.11.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.14.norm2.weight', 'model.visual_model.image_encoder.blocks.21.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.28.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.13.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.24.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.15.attn.proj.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.bias', 'model.visual_model.image_encoder.blocks.13.norm2.bias', 'model.visual_model.image_encoder.blocks.14.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.7.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.5.norm2.bias', 'model.visual_model.mask_decoder.transformer.layers.1.norm1.weight', 'model.visual_model.image_encoder.blocks.19.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.6.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.20.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.19.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.26.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.7.attn.proj.bias', 'mask_pooler.block1.0.0.bias', 'model.visual_model.mask_decoder.iou_prediction_head.layers.2.bias', 'model.visual_model.image_encoder.blocks.20.attn.proj.weight', 'model.visual_model.prompt_encoder.mask_downscaling.6.bias', 'model.visual_model.image_encoder.blocks.22.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.5.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.27.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.29.mlp.lin1.weight', 'bio_encoder.encoder.layers.0.norm2.weight', 'model.visual_model.image_encoder.blocks.2.attn.proj.weight', 'model.visual_model.image_encoder.blocks.27.norm1.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.weight', 'model.visual_model.mask_decoder.iou_prediction_head.layers.1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.weight', 'model.visual_model.image_encoder.blocks.11.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.9.norm1.weight', 'model.visual_model.image_encoder.blocks.10.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.20.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.6.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.6.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.22.norm2.bias', 'mask_pooler.block2.0.4.bias', 'mask_pooler.prompt_encoder.mask_downscaling.1.bias', 'model.text_hidden_fcs.0.2.bias', 'model.visual_model.image_encoder.blocks.2.norm2.weight', 'model.visual_model.image_encoder.blocks.10.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.20.norm1.bias', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.bias', 'model.visual_model.image_encoder.blocks.19.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.5.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.1.norm1.bias', 'mask_pooler.block1.0.0.weight', 'model.visual_model.image_encoder.blocks.26.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.0.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.14.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.18.norm2.bias', 'bio_encoder.encoder.layers.0.self_attn.out_proj.weight', 'model.visual_model.image_encoder.blocks.19.norm1.weight', 'model.visual_model.image_encoder.blocks.6.attn.proj.weight', 'model.visual_model.image_encoder.blocks.29.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.norm2.weight', 'model.visual_model.image_encoder.neck.2.weight', 'model.visual_model.image_encoder.blocks.24.attn.qkv.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm2.weight', 'model.visual_model.image_encoder.blocks.2.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.4.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.19.attn.proj.weight', 'model.visual_model.image_encoder.blocks.3.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.bias', 'mask_pooler.prompt_encoder.mask_downscaling.0.bias', 'model.visual_model.image_encoder.blocks.6.attn.rel_pos_h', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.bias', 'model.visual_model.image_encoder.blocks.26.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.weight', 'model.visual_model.image_encoder.blocks.5.norm1.weight', 'model.visual_model.image_encoder.blocks.20.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.12.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.29.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.20.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.5.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.17.attn.qkv.bias', 'mask_pooler.mask_modulator.projections.2.bias', 'model.visual_model.image_encoder.blocks.4.norm1.weight', 'model.visual_model.image_encoder.blocks.11.norm1.bias', 'model.visual_model.mask_decoder.output_upscaling.1.weight', 'model.visual_model.image_encoder.blocks.19.attn.rel_pos_w', 'mask_pooler.prompt_encoder.not_a_point_embed.weight', 'model.visual_model.image_encoder.blocks.14.attn.rel_pos_w', 'model.visual_model.mask_decoder.transformer.layers.1.norm3.bias', 'model.visual_model.image_encoder.blocks.12.attn.proj.weight', 'model.visual_model.image_encoder.blocks.8.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.6.attn.proj.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.weight', 'model.visual_model.image_encoder.blocks.25.norm2.bias', 'model.visual_model.image_encoder.blocks.30.attn.proj.weight', 'model.visual_model.image_encoder.blocks.26.norm2.bias', 'model.visual_model.prompt_encoder.point_embeddings.1.weight', 'model.visual_model.image_encoder.blocks.9.norm2.weight', 'model.visual_model.image_encoder.blocks.16.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.28.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.29.attn.proj.bias', 'model.visual_model.image_encoder.blocks.29.attn.rel_pos_w', 'mask_pooler.prompt_encoder.point_embeddings.0.weight', 'model.visual_model.image_encoder.blocks.15.norm2.bias', 'model.visual_model.image_encoder.blocks.23.norm1.weight', 'model.token_to_mask_fcs.0.0.bias', 'model.visual_model.image_encoder.blocks.13.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.10.norm2.bias', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.bias', 'model.visual_model.image_encoder.blocks.29.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.29.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.3.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.27.norm2.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.weight', 'model.visual_model.image_encoder.blocks.11.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.24.norm2.bias', 'model.visual_model.image_encoder.blocks.22.mlp.lin2.bias', 'model.text_hidden_fcs.0.0.bias', 'mask_pooler.mask_modulator.projections.0.weight', 'model.visual_model.image_encoder.blocks.30.norm2.weight', 'model.visual_model.image_encoder.blocks.18.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.15.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.28.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.25.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.4.attn.proj.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.weight', 'model.visual_model.image_encoder.blocks.17.attn.proj.bias', 'model.visual_model.image_encoder.blocks.0.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.1.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.16.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.19.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.bias', 'model.visual_model.image_encoder.blocks.13.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.6.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.18.attn.proj.weight', 'model.visual_model.image_encoder.blocks.23.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.9.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.24.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.24.attn.proj.bias', 'model.visual_model.image_encoder.blocks.10.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.15.attn.qkv.weight', 'model.token_to_mask_fcs.0.0.weight', 'mask_pooler.block1.0.1.bias', 'model.visual_model.image_encoder.blocks.12.norm2.bias', 'model.visual_model.image_encoder.blocks.1.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.7.mlp.lin1.weight', 'model.visual_model.mask_decoder.mask_tokens.weight', 'model.visual_model.image_encoder.blocks.11.attn.rel_pos_w', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.bias', 'model.visual_model.image_encoder.blocks.17.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.23.attn.proj.weight', 'model.visual_model.image_encoder.blocks.9.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.29.norm1.weight', 'model.visual_model.image_encoder.blocks.29.norm2.bias', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.bias', 'model.visual_model.prompt_encoder.mask_downscaling.6.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.bias', 'mask_pooler.mask_modulator.projections.1.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.bias', 'model.visual_model.image_encoder.blocks.10.attn.proj.bias', 'model.visual_model.image_encoder.blocks.8.attn.proj.weight', 'model.visual_model.image_encoder.blocks.29.attn.proj.weight', 'bio_encoder.encoder.layers.0.linear2.bias', 'mask_pooler.attn.out_proj.weight', 'model.visual_model.image_encoder.blocks.19.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.20.norm1.weight', 'model.visual_model.image_encoder.blocks.31.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.30.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.28.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.3.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.weight', 'model.visual_model.image_encoder.blocks.7.attn.proj.weight', 'model.visual_model.image_encoder.blocks.31.norm1.bias', 'model.visual_model.image_encoder.blocks.12.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.30.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.0.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.27.norm1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.bias', 'model.visual_model.image_encoder.blocks.17.attn.proj.weight', 'model.visual_model.image_encoder.blocks.7.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.weight', 'model.visual_model.image_encoder.blocks.5.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.3.norm2.weight', 'model.visual_model.image_encoder.blocks.30.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.2.mlp.lin2.bias', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.bias', 'model.visual_model.image_encoder.blocks.5.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.31.norm2.weight', 'model.visual_model.image_encoder.blocks.27.attn.proj.weight', 'model.visual_model.image_encoder.blocks.27.mlp.lin2.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.weight', 'model.visual_model.image_encoder.blocks.1.attn.proj.bias', 'model.visual_model.image_encoder.blocks.5.attn.proj.bias', 'model.visual_model.image_encoder.blocks.29.norm2.weight', 'model.visual_model.image_encoder.blocks.28.attn.proj.weight', 'model.visual_model.image_encoder.blocks.27.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.4.norm1.bias', 'model.visual_model.image_encoder.blocks.0.attn.proj.weight', 'model.visual_model.image_encoder.blocks.16.norm2.bias', 'model.visual_model.image_encoder.blocks.23.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.17.mlp.lin2.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.bias', 'model.visual_model.image_encoder.blocks.23.norm1.bias', 'model.visual_model.image_encoder.blocks.8.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.14.attn.rel_pos_h', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.weight', 'model.visual_model.image_encoder.blocks.22.attn.rel_pos_h', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.bias', 'bio_encoder.encoder.layers.0.linear2.weight', 'model.visual_model.image_encoder.blocks.24.norm1.weight', 'model.visual_model.image_encoder.blocks.29.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.24.norm1.bias', 'model.visual_model.mask_decoder.iou_prediction_head.layers.0.weight', 'model.visual_model.image_encoder.blocks.23.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.16.norm1.bias', 'model.visual_model.image_encoder.blocks.24.norm2.weight', 'model.visual_model.image_encoder.blocks.12.mlp.lin1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.weight', 'model.visual_model.image_encoder.blocks.28.norm2.bias', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.bias', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.weight', 'model.visual_model.image_encoder.blocks.27.attn.proj.bias', 'model.visual_model.image_encoder.blocks.28.norm2.weight', 'model.visual_model.image_encoder.blocks.23.norm2.weight', 'model.visual_model.image_encoder.blocks.31.norm2.bias', 'model.visual_model.image_encoder.blocks.2.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.30.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.24.attn.proj.weight', 'model.visual_model.image_encoder.blocks.6.norm2.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.bias', 'bio_encoder.encoder.layers.0.self_attn.out_proj.bias', 'bio_encoder.encoder.layers.0.self_attn.in_proj_weight', 'model.visual_model.image_encoder.blocks.12.norm1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.weight', 'model.visual_model.image_encoder.blocks.11.mlp.lin1.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.weight', 'model.visual_model.image_encoder.blocks.13.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.weight', 'model.visual_model.image_encoder.blocks.16.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.23.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.17.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm2.bias', 'model.visual_model.image_encoder.blocks.23.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.2.norm2.bias', 'model.visual_model.image_encoder.blocks.17.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.12.attn.proj.bias', 'model.visual_model.image_encoder.blocks.17.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.30.norm2.bias', 'model.visual_model.image_encoder.blocks.9.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.8.norm1.bias', 'model.visual_model.image_encoder.blocks.28.attn.proj.bias', 'model.visual_model.image_encoder.blocks.3.attn.proj.weight', 'model.visual_model.mask_decoder.iou_token.weight', 'model.visual_model.mask_decoder.output_upscaling.0.weight', 'model.visual_model.image_encoder.blocks.13.attn.proj.weight', 'model.visual_model.image_encoder.blocks.18.attn.proj.bias', 'model.visual_model.image_encoder.blocks.16.attn.proj.weight', 'model.visual_model.image_encoder.blocks.27.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.25.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.12.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.1.norm1.weight', 'model.visual_model.image_encoder.blocks.9.mlp.lin1.weight', 'model.visual_model.prompt_encoder.no_mask_embed.weight', 'model.visual_model.image_encoder.blocks.2.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.norm2.bias', 'model.visual_model.image_encoder.blocks.2.attn.proj.bias', 'model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.7.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.3.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.3.attn.proj.bias', 'model.visual_model.image_encoder.blocks.30.norm1.weight', 'model.visual_model.image_encoder.blocks.4.mlp.lin1.weight', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.weight', 'model.visual_model.image_encoder.blocks.6.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.26.attn.proj.bias', 'model.visual_model.image_encoder.blocks.2.norm1.weight', 'model.visual_model.image_encoder.blocks.16.norm1.weight', 'bio_encoder.encoder.layers.0.norm1.bias', 'model.visual_model.image_encoder.blocks.7.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.2.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.8.attn.rel_pos_h', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.weight', 'model.visual_model.image_encoder.blocks.22.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.18.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.18.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.25.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.weight', 'model.visual_model.image_encoder.blocks.9.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.bias', 'model.visual_model.image_encoder.blocks.9.mlp.lin1.bias', 'model.visual_model.image_encoder.neck.1.bias', 'model.visual_model.image_encoder.blocks.1.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.14.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.10.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.2.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.4.norm2.bias', 'model.visual_model.image_encoder.blocks.16.attn.rel_pos_h', 'model.visual_model.mask_decoder.iou_prediction_head.layers.0.bias', 'model.visual_model.image_encoder.blocks.29.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.15.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.14.mlp.lin2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.bias', 'model.visual_model.image_encoder.blocks.25.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.28.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.23.attn.proj.bias', 'model.visual_model.image_encoder.blocks.4.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.11.attn.proj.weight', 'model.visual_model.image_encoder.blocks.8.norm1.weight', 'model.visual_model.image_encoder.blocks.31.norm1.weight', 'model.visual_model.prompt_encoder.mask_downscaling.0.weight', 'mask_pooler.mask_modulator.projections.0.bias', 'mask_pooler.attn.in_proj_weight', 'model.visual_model.image_encoder.blocks.28.attn.qkv.bias', 'mask_pooler.queries', 'model.visual_model.image_encoder.blocks.10.norm2.weight', 'mask_pooler.block2.0.4.weight', 'model.visual_model.image_encoder.blocks.4.norm2.weight', 'model.visual_model.mask_decoder.transformer.layers.1.norm4.weight', 'model.visual_model.image_encoder.blocks.15.mlp.lin1.weight', 'mask_pooler.prompt_encoder.mask_downscaling.3.bias', 'model.visual_model.image_encoder.blocks.21.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.7.attn.qkv.bias', 'model.visual_model.prompt_encoder.point_embeddings.0.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.weight', 'model.visual_model.image_encoder.blocks.7.norm1.weight', 'model.visual_model.image_encoder.blocks.11.norm2.weight', 'model.token_to_mask_fcs.0.2.weight', 'model.visual_model.image_encoder.blocks.28.norm1.bias', 'model.visual_model.image_encoder.blocks.28.mlp.lin2.weight', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.weight', 'model.visual_model.image_encoder.blocks.19.mlp.lin2.bias', 'mask_pooler.attn.in_proj_bias', 'model.visual_model.image_encoder.blocks.21.attn.qkv.bias', 'model.visual_model.prompt_encoder.not_a_point_embed.weight', 'model.token_to_mask_fcs.0.2.bias', 'model.visual_model.prompt_encoder.mask_downscaling.3.weight', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.weight', 'bio_encoder.encoder.layers.0.linear1.weight', 'model.visual_model.image_encoder.blocks.13.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.15.attn.proj.bias', 'mask_pooler.mask_modulator.projections.1.weight', 'model.visual_model.image_encoder.blocks.10.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.weight', 'model.visual_model.image_encoder.blocks.28.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.23.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.15.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.27.norm2.weight', 'mask_pooler.prompt_encoder.point_embeddings.2.weight', 'mask_pooler.attn.out_proj.bias', 'model.visual_model.image_encoder.blocks.25.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.30.attn.proj.bias', 'model.visual_model.image_encoder.blocks.31.attn.rel_pos_h', 'bio_encoder.encoder.layers.0.self_attn.in_proj_bias', 'model.visual_model.image_encoder.blocks.0.attn.proj.bias', 'model.visual_model.image_encoder.blocks.4.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.24.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.20.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.18.mlp.lin1.weight', 'mask_pooler.prompt_encoder.mask_downscaling.6.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm1.bias', 'model.visual_model.image_encoder.blocks.22.norm1.bias', 'model.visual_model.image_encoder.blocks.0.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.6.norm2.weight', 'model.visual_model.image_encoder.blocks.25.norm1.weight', 'mask_pooler.prompt_encoder.mask_downscaling.4.weight', 'model.visual_model.image_encoder.blocks.9.attn.proj.weight', 'model.visual_model.image_encoder.blocks.8.norm2.weight', 'model.visual_model.image_encoder.blocks.13.attn.proj.bias', 'model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.5.norm1.bias', 'model.visual_model.image_encoder.blocks.14.attn.proj.weight', 'mask_pooler.prompt_encoder.mask_downscaling.6.bias', 'model.visual_model.mask_decoder.output_upscaling.3.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm3.bias', 'model.visual_model.image_encoder.blocks.10.attn.proj.weight', 'model.visual_model.mask_decoder.transformer.norm_final_attn.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.bias', 'model.visual_model.mask_decoder.iou_prediction_head.layers.1.weight', 'model.visual_model.image_encoder.blocks.7.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.7.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.23.mlp.lin2.weight', 'model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.norm4.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.weight', 'model.visual_model.image_encoder.blocks.21.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.25.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.0.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.11.norm2.bias', 'model.visual_model.image_encoder.blocks.18.attn.qkv.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.weight', 'model.visual_model.image_encoder.blocks.6.norm1.bias', 'model.visual_model.image_encoder.blocks.13.norm1.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.bias', 'model.visual_model.image_encoder.blocks.26.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.3.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.22.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.30.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.5.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.21.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.9.attn.qkv.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.weight', 'model.visual_model.image_encoder.blocks.1.norm2.bias', 'model.visual_model.image_encoder.blocks.16.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.19.norm2.weight', 'model.visual_model.image_encoder.blocks.11.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.22.attn.rel_pos_w', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.weight', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.bias', 'model.visual_model.image_encoder.patch_embed.proj.weight', 'model.visual_model.image_encoder.blocks.19.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.21.attn.proj.weight', 'model.visual_model.image_encoder.blocks.15.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.0.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.25.attn.proj.bias', 'model.visual_model.mask_decoder.output_upscaling.0.bias', 'model.visual_model.image_encoder.blocks.6.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.21.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.1.attn.qkv.weight', 'model.visual_model.image_encoder.blocks.13.attn.qkv.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.weight', 'model.text_hidden_fcs.0.2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.norm4.bias', 'model.visual_model.image_encoder.blocks.12.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.10.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.19.norm2.bias', 'model.visual_model.prompt_encoder.point_embeddings.2.weight', 'model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.4.attn.rel_pos_w', 'model.visual_model.prompt_encoder.mask_downscaling.1.bias', 'model.visual_model.image_encoder.blocks.3.attn.qkv.bias', 'model.visual_model.prompt_encoder.point_embeddings.3.weight', 'model.visual_model.image_encoder.blocks.20.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.21.norm1.bias', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.weight', 'model.visual_model.image_encoder.blocks.22.attn.proj.weight', 'model.visual_model.image_encoder.blocks.12.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.31.attn.proj.weight', 'model.visual_model.image_encoder.blocks.11.attn.proj.bias', 'model.visual_model.image_encoder.blocks.6.norm1.weight', 'model.visual_model.image_encoder.blocks.3.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.25.norm1.bias', 'mask_pooler.prompt_encoder.mask_downscaling.0.weight', 'model.visual_model.image_encoder.blocks.0.attn.rel_pos_w', 'model.visual_model.prompt_encoder.mask_downscaling.4.bias', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.bias', 'model.visual_model.image_encoder.blocks.8.norm2.bias', 'model.visual_model.image_encoder.blocks.26.attn.proj.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.bias', 'model.visual_model.image_encoder.blocks.3.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.11.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.21.attn.proj.bias', 'model.visual_model.prompt_encoder.mask_downscaling.1.weight', 'model.visual_model.image_encoder.blocks.5.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.1.norm1.bias', 'model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.weight', 'model.visual_model.image_encoder.blocks.14.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.14.attn.proj.bias', 'model.visual_model.image_encoder.blocks.18.norm1.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.bias', 'model.visual_model.image_encoder.blocks.13.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.14.norm2.bias', 'model.visual_model.image_encoder.blocks.23.norm2.bias', 'model.visual_model.image_encoder.blocks.16.attn.proj.bias', 'model.visual_model.image_encoder.blocks.31.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.16.norm2.weight', 'model.visual_model.mask_decoder.transformer.norm_final_attn.weight', 'model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.bias', 'model.visual_model.image_encoder.blocks.14.norm1.bias', 'bio_encoder.encoder.layers.0.norm2.bias', 'model.visual_model.image_encoder.blocks.26.norm1.weight', 'model.visual_model.image_encoder.blocks.7.norm2.bias', 'model.visual_model.image_encoder.blocks.21.norm1.weight', 'bio_encoder.encoder.layers.0.norm1.weight', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.weight', 'model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.bias', 'model.visual_model.mask_decoder.transformer.layers.0.norm1.weight', 'model.visual_model.prompt_encoder.mask_downscaling.3.bias', 'model.visual_model.image_encoder.blocks.9.norm2.bias', 'model.visual_model.image_encoder.blocks.8.mlp.lin1.weight', 'model.visual_model.image_encoder.blocks.21.norm2.bias', 'model.visual_model.image_encoder.blocks.25.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.10.attn.qkv.bias', 'model.visual_model.image_encoder.blocks.30.mlp.lin2.weight', 'mask_pooler.prompt_encoder.mask_downscaling.3.weight', 'model.visual_model.image_encoder.blocks.8.mlp.lin2.weight', 'mask_pooler.block1.0.1.weight', 'model.visual_model.image_encoder.blocks.20.attn.rel_pos_w', 'model.visual_model.image_encoder.blocks.15.norm2.weight', 'mask_pooler.final_conv.bias', 'model.visual_model.mask_decoder.transformer.layers.1.norm3.weight', 'model.visual_model.image_encoder.blocks.8.attn.proj.bias', 'model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.weight', 'model.visual_model.image_encoder.blocks.13.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.26.norm1.bias', 'model.visual_model.image_encoder.blocks.26.mlp.lin2.bias', 'model.visual_model.image_encoder.blocks.24.mlp.lin1.bias', 'model.visual_model.image_encoder.blocks.15.norm1.weight', 'model.visual_model.image_encoder.blocks.29.attn.rel_pos_h', 'model.visual_model.image_encoder.blocks.9.mlp.lin2.weight', 'model.visual_model.image_encoder.blocks.25.mlp.lin2.bias', 'model.visual_model.prompt_encoder.mask_downscaling.4.weight', 'model.visual_model.image_encoder.blocks.22.norm2.weight', 'model.visual_model.image_encoder.blocks.31.mlp.lin2.weight', 'model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.weight', 'model.visual_model.image_encoder.blocks.26.attn.rel_pos_w']
You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.
trainable params: 6,553,600 || all params: 14,153,038,759 || trainable%: 0.04630525014165263
trainable params: 6,553,600 || all params: 14,153,038,759 || trainable%: 0.04630525014165263
>> model.config.train_mask_prompt_encoder:  True
n:  base_model.model.model.embed_tokens.weight p.shape:  torch.Size([32002, 5120])
n:  base_model.model.model.visual_model.prompt_encoder.point_embeddings.0.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.point_embeddings.1.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.point_embeddings.2.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.point_embeddings.3.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.not_a_point_embed.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.0.weight p.shape:  torch.Size([4, 1, 2, 2])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.0.bias p.shape:  torch.Size([4])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.1.weight p.shape:  torch.Size([4])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.1.bias p.shape:  torch.Size([4])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.3.weight p.shape:  torch.Size([16, 4, 2, 2])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.3.bias p.shape:  torch.Size([16])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.4.weight p.shape:  torch.Size([16])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.4.bias p.shape:  torch.Size([16])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.6.weight p.shape:  torch.Size([256, 16, 1, 1])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.6.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.prompt_encoder.no_mask_embed.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm1.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm2.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.weight p.shape:  torch.Size([2048, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.bias p.shape:  torch.Size([2048])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.weight p.shape:  torch.Size([256, 2048])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm3.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm3.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm4.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm4.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm1.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm2.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.weight p.shape:  torch.Size([2048, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.bias p.shape:  torch.Size([2048])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.weight p.shape:  torch.Size([256, 2048])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm3.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm3.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm4.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm4.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.norm_final_attn.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.norm_final_attn.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.iou_token.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.mask_decoder.mask_tokens.weight p.shape:  torch.Size([4, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.0.weight p.shape:  torch.Size([256, 64, 2, 2])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.0.bias p.shape:  torch.Size([64])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.1.weight p.shape:  torch.Size([64])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.1.bias p.shape:  torch.Size([64])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.3.weight p.shape:  torch.Size([64, 32, 2, 2])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.3.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.2.weight p.shape:  torch.Size([4, 256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.2.bias p.shape:  torch.Size([4])
n:  base_model.model.model.text_hidden_fcs.0.0.weight p.shape:  torch.Size([5120, 5120])
n:  base_model.model.model.text_hidden_fcs.0.0.bias p.shape:  torch.Size([5120])
n:  base_model.model.model.text_hidden_fcs.0.2.weight p.shape:  torch.Size([256, 5120])
n:  base_model.model.model.text_hidden_fcs.0.2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.token_to_mask_fcs.0.0.weight p.shape:  torch.Size([5120, 5120])
n:  base_model.model.model.token_to_mask_fcs.0.0.bias p.shape:  torch.Size([5120])
n:  base_model.model.model.token_to_mask_fcs.0.2.weight p.shape:  torch.Size([3, 5120])
n:  base_model.model.model.token_to_mask_fcs.0.2.bias p.shape:  torch.Size([3])
n:  base_model.model.lm_head.weight p.shape:  torch.Size([32002, 5120])
n:  base_model.model.bio_encoder.encoder.layers.0.self_attn.in_proj_weight p.shape:  torch.Size([15360, 5120])
n:  base_model.model.bio_encoder.encoder.layers.0.self_attn.in_proj_bias p.shape:  torch.Size([15360])
n:  base_model.model.bio_encoder.encoder.layers.0.self_attn.out_proj.weight p.shape:  torch.Size([5120, 5120])
n:  base_model.model.bio_encoder.encoder.layers.0.self_attn.out_proj.bias p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.linear1.weight p.shape:  torch.Size([2048, 5120])
n:  base_model.model.bio_encoder.encoder.layers.0.linear1.bias p.shape:  torch.Size([2048])
n:  base_model.model.bio_encoder.encoder.layers.0.linear2.weight p.shape:  torch.Size([5120, 2048])
n:  base_model.model.bio_encoder.encoder.layers.0.linear2.bias p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.norm1.weight p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.norm1.bias p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.norm2.weight p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.norm2.bias p.shape:  torch.Size([5120])
n:  base_model.model.mask_pooler.queries p.shape:  torch.Size([4096, 1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.point_embeddings.0.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.point_embeddings.1.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.point_embeddings.2.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.point_embeddings.3.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.not_a_point_embed.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.0.weight p.shape:  torch.Size([4, 1, 2, 2])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.0.bias p.shape:  torch.Size([4])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.1.weight p.shape:  torch.Size([4])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.1.bias p.shape:  torch.Size([4])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.3.weight p.shape:  torch.Size([16, 4, 2, 2])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.3.bias p.shape:  torch.Size([16])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.4.weight p.shape:  torch.Size([16])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.4.bias p.shape:  torch.Size([16])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.6.weight p.shape:  torch.Size([256, 16, 1, 1])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.6.bias p.shape:  torch.Size([256])
n:  base_model.model.mask_pooler.prompt_encoder.no_mask_embed.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.mask_modulator.projections.0.weight p.shape:  torch.Size([8, 256])
n:  base_model.model.mask_pooler.mask_modulator.projections.0.bias p.shape:  torch.Size([8])
n:  base_model.model.mask_pooler.mask_modulator.projections.1.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.mask_pooler.mask_modulator.projections.1.bias p.shape:  torch.Size([32])
n:  base_model.model.mask_pooler.mask_modulator.projections.2.weight p.shape:  torch.Size([512, 256])
n:  base_model.model.mask_pooler.mask_modulator.projections.2.bias p.shape:  torch.Size([512])
n:  base_model.model.mask_pooler.attn.in_proj_weight p.shape:  torch.Size([768, 256])
n:  base_model.model.mask_pooler.attn.in_proj_bias p.shape:  torch.Size([768])
n:  base_model.model.mask_pooler.attn.out_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.mask_pooler.attn.out_proj.bias p.shape:  torch.Size([256])
ade20k:  20210
cocostuff:  118287
loading annotations into memory...
Done (t=0.62s)
creating index...
index created!
pascal_part:  4366
loading annotations into memory...
>> model.config.train_mask_prompt_encoder:  True
n:  base_model.model.model.embed_tokens.weight p.shape:  torch.Size([32002, 5120])
n:  base_model.model.model.visual_model.prompt_encoder.point_embeddings.0.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.point_embeddings.1.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.point_embeddings.2.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.point_embeddings.3.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.not_a_point_embed.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.0.weight p.shape:  torch.Size([4, 1, 2, 2])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.0.bias p.shape:  torch.Size([4])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.1.weight p.shape:  torch.Size([4])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.1.bias p.shape:  torch.Size([4])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.3.weight p.shape:  torch.Size([16, 4, 2, 2])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.3.bias p.shape:  torch.Size([16])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.4.weight p.shape:  torch.Size([16])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.4.bias p.shape:  torch.Size([16])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.6.weight p.shape:  torch.Size([256, 16, 1, 1])
n:  base_model.model.model.visual_model.prompt_encoder.mask_downscaling.6.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.prompt_encoder.no_mask_embed.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm1.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm2.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.weight p.shape:  torch.Size([2048, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.bias p.shape:  torch.Size([2048])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.weight p.shape:  torch.Size([256, 2048])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm3.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm3.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm4.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.norm4.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm1.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm2.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.weight p.shape:  torch.Size([2048, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.bias p.shape:  torch.Size([2048])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.weight p.shape:  torch.Size([256, 2048])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm3.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm3.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm4.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.norm4.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.weight p.shape:  torch.Size([128, 256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.bias p.shape:  torch.Size([128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.weight p.shape:  torch.Size([256, 128])
n:  base_model.model.model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.norm_final_attn.weight p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.transformer.norm_final_attn.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.iou_token.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.model.visual_model.mask_decoder.mask_tokens.weight p.shape:  torch.Size([4, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.0.weight p.shape:  torch.Size([256, 64, 2, 2])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.0.bias p.shape:  torch.Size([64])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.1.weight p.shape:  torch.Size([64])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.1.bias p.shape:  torch.Size([64])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.3.weight p.shape:  torch.Size([64, 32, 2, 2])
n:  base_model.model.model.visual_model.mask_decoder.output_upscaling.3.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.bias p.shape:  torch.Size([32])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.0.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.0.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.1.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.1.bias p.shape:  torch.Size([256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.2.weight p.shape:  torch.Size([4, 256])
n:  base_model.model.model.visual_model.mask_decoder.iou_prediction_head.layers.2.bias p.shape:  torch.Size([4])
n:  base_model.model.model.text_hidden_fcs.0.0.weight p.shape:  torch.Size([5120, 5120])
n:  base_model.model.model.text_hidden_fcs.0.0.bias p.shape:  torch.Size([5120])
n:  base_model.model.model.text_hidden_fcs.0.2.weight p.shape:  torch.Size([256, 5120])
n:  base_model.model.model.text_hidden_fcs.0.2.bias p.shape:  torch.Size([256])
n:  base_model.model.model.token_to_mask_fcs.0.0.weight p.shape:  torch.Size([5120, 5120])
n:  base_model.model.model.token_to_mask_fcs.0.0.bias p.shape:  torch.Size([5120])
n:  base_model.model.model.token_to_mask_fcs.0.2.weight p.shape:  torch.Size([3, 5120])
n:  base_model.model.model.token_to_mask_fcs.0.2.bias p.shape:  torch.Size([3])
n:  base_model.model.lm_head.weight p.shape:  torch.Size([32002, 5120])
n:  base_model.model.bio_encoder.encoder.layers.0.self_attn.in_proj_weight p.shape:  torch.Size([15360, 5120])
n:  base_model.model.bio_encoder.encoder.layers.0.self_attn.in_proj_bias p.shape:  torch.Size([15360])
n:  base_model.model.bio_encoder.encoder.layers.0.self_attn.out_proj.weight p.shape:  torch.Size([5120, 5120])
n:  base_model.model.bio_encoder.encoder.layers.0.self_attn.out_proj.bias p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.linear1.weight p.shape:  torch.Size([2048, 5120])
n:  base_model.model.bio_encoder.encoder.layers.0.linear1.bias p.shape:  torch.Size([2048])
n:  base_model.model.bio_encoder.encoder.layers.0.linear2.weight p.shape:  torch.Size([5120, 2048])
n:  base_model.model.bio_encoder.encoder.layers.0.linear2.bias p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.norm1.weight p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.norm1.bias p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.norm2.weight p.shape:  torch.Size([5120])
n:  base_model.model.bio_encoder.encoder.layers.0.norm2.bias p.shape:  torch.Size([5120])
n:  base_model.model.mask_pooler.queries p.shape:  torch.Size([4096, 1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.point_embeddings.0.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.point_embeddings.1.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.point_embeddings.2.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.point_embeddings.3.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.not_a_point_embed.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.0.weight p.shape:  torch.Size([4, 1, 2, 2])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.0.bias p.shape:  torch.Size([4])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.1.weight p.shape:  torch.Size([4])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.1.bias p.shape:  torch.Size([4])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.3.weight p.shape:  torch.Size([16, 4, 2, 2])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.3.bias p.shape:  torch.Size([16])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.4.weight p.shape:  torch.Size([16])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.4.bias p.shape:  torch.Size([16])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.6.weight p.shape:  torch.Size([256, 16, 1, 1])
n:  base_model.model.mask_pooler.prompt_encoder.mask_downscaling.6.bias p.shape:  torch.Size([256])
n:  base_model.model.mask_pooler.prompt_encoder.no_mask_embed.weight p.shape:  torch.Size([1, 256])
n:  base_model.model.mask_pooler.mask_modulator.projections.0.weight p.shape:  torch.Size([8, 256])
n:  base_model.model.mask_pooler.mask_modulator.projections.0.bias p.shape:  torch.Size([8])
n:  base_model.model.mask_pooler.mask_modulator.projections.1.weight p.shape:  torch.Size([32, 256])
n:  base_model.model.mask_pooler.mask_modulator.projections.1.bias p.shape:  torch.Size([32])
n:  base_model.model.mask_pooler.mask_modulator.projections.2.weight p.shape:  torch.Size([512, 256])
n:  base_model.model.mask_pooler.mask_modulator.projections.2.bias p.shape:  torch.Size([512])
n:  base_model.model.mask_pooler.attn.in_proj_weight p.shape:  torch.Size([768, 256])
n:  base_model.model.mask_pooler.attn.in_proj_bias p.shape:  torch.Size([768])
n:  base_model.model.mask_pooler.attn.out_proj.weight p.shape:  torch.Size([256, 256])
n:  base_model.model.mask_pooler.attn.out_proj.bias p.shape:  torch.Size([256])
ade20k:  20210
cocostuff:  118287
loading annotations into memory...
Done (t=0.54s)
creating index...
index created!
pascal_part:  4366
loading annotations into memory...
Done (t=8.52s)
creating index...
index created!
paco_lvis:  45790
mapillary:  18000
loading dataset refclef into memory...
ref_file:  /shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/refclef/refs(unc).p
creating index...
index created.
DONE (t=2.98s)
dataset refclef (refs unc) (train split) has 17978 images and 99523 annotations.
loading dataset refcoco into memory...
ref_file:  /shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/refcoco/refs(unc).p
creating index...
Done (t=8.39s)
creating index...
index created!
paco_lvis:  45790
mapillary:  18000
loading dataset refclef into memory...
ref_file:  /shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/refclef/refs(unc).p
index created.
DONE (t=3.57s)
dataset refcoco (refs unc) (train split) has 16994 images and 196771 annotations.
loading dataset refcoco+ into memory...
ref_file:  /shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/refcoco+/refs(unc).p
creating index...
index created.
DONE (t=3.00s)
dataset refclef (refs unc) (train split) has 17978 images and 99523 annotations.
loading dataset refcoco into memory...
ref_file:  /shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/refcoco/refs(unc).p
creating index...
creating index...
index created.
DONE (t=3.60s)
dataset refcoco (refs unc) (train split) has 16994 images and 196771 annotations.
loading dataset refcoco+ into memory...
ref_file:  /shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/refcoco+/refs(unc).p
index created.
DONE (t=6.16s)
dataset refcoco+ (refs unc) (train split) has 16992 images and 196737 annotations.
loading dataset refcocog into memory...
ref_file:  /shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/refcocog/refs(umd).p
creating index...
creating index...
index created.
DONE (t=5.78s)
dataset refcoco+ (refs unc) (train split) has 16992 images and 196737 annotations.
loading dataset refcocog into memory...
ref_file:  /shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/refcocog/refs(umd).p
index created.
DONE (t=6.35s)
dataset refcocog (refs umd) (train split) has 21899 images and 208960 annotations.
llava_instruct_150k_data:  157712
number of reason_seg samples:  239
len(self.img_to_explanation):  239
Training with 60000 examples.
Validating with 200 examples.
[2025-10-22 16:10:49,828] [INFO] [logging.py:96:log_dist] [Rank -1] DeepSpeed info: version=0.9.5, git-hash=unknown, git-branch=unknown
[2025-10-22 16:10:49,829] [WARNING] [comm.py:152:init_deepspeed_backend] NCCL backend in DeepSpeed not yet implemented
[2025-10-22 16:10:49,829] [INFO] [comm.py:594:init_distributed] cdb=None
creating index...
index created.
DONE (t=6.03s)
dataset refcocog (refs umd) (train split) has 21899 images and 208960 annotations.
llava_instruct_150k_data:  157712
number of reason_seg samples:  239
len(self.img_to_explanation):  239
Training with 60000 examples.
Validating with 200 examples.
[2025-10-22 16:10:55,033] [INFO] [logging.py:96:log_dist] [Rank -1] DeepSpeed info: version=0.9.5, git-hash=unknown, git-branch=unknown
[2025-10-22 16:10:55,033] [WARNING] [comm.py:152:init_deepspeed_backend] NCCL backend in DeepSpeed not yet implemented
[2025-10-22 16:10:55,033] [INFO] [comm.py:594:init_distributed] cdb=None
[2025-10-22 16:10:55,033] [INFO] [comm.py:625:init_distributed] Initializing TorchBackend in DeepSpeed with backend nccl
[2025-10-22 16:11:07,700] [INFO] [logging.py:96:log_dist] [Rank 0] DeepSpeed Flops Profiler Enabled: False
Using /shared/nas/data/m1/jk100/.cache/torch_extensions/py310_cu118 as PyTorch extensions root...Using /shared/nas/data/m1/jk100/.cache/torch_extensions/py310_cu118 as PyTorch extensions root...

Detected CUDA files, patching ldflags
Emitting ninja build file /shared/nas/data/m1/jk100/.cache/torch_extensions/py310_cu118/fused_adam/build.ninja...
/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/cpp_extension.py:1967: UserWarning: TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation. 
If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'].
  warnings.warn(
Building extension module fused_adam...
Allowing ninja to set a default number of workers... (overridable by setting the environment variable MAX_JOBS=N)
ninja: no work to do.
Loading extension module fused_adam...
Time to load fused_adam op: 0.5220873355865479 seconds
/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/ops/adam/fused_adam.py:96: UserWarning: The torch.cuda.*DtypeTensor constructors are no longer recommended. It's best to use methods such as torch.tensor(data, dtype=*, device='cuda') to create tensors. (Triggered internally at /opt/conda/conda-bld/pytorch_1712608839953/work/torch/csrc/tensor/python_tensor.cpp:78.)
  self._dummy_overflow_buf = get_accelerator().IntTensor([0])
Loading extension module fused_adam...
Time to load fused_adam op: 0.5270330905914307 seconds
/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/ops/adam/fused_adam.py:96: UserWarning: The torch.cuda.*DtypeTensor constructors are no longer recommended. It's best to use methods such as torch.tensor(data, dtype=*, device='cuda') to create tensors. (Triggered internally at /opt/conda/conda-bld/pytorch_1712608839953/work/torch/csrc/tensor/python_tensor.cpp:78.)
  self._dummy_overflow_buf = get_accelerator().IntTensor([0])
[2025-10-22 16:11:08,411] [INFO] [logging.py:96:log_dist] [Rank 0] Using DeepSpeed Optimizer param name adamw as basic optimizer
[2025-10-22 16:11:08,601] [INFO] [logging.py:96:log_dist] [Rank 0] DeepSpeed Basic Optimizer = FusedAdam
[2025-10-22 16:11:08,602] [INFO] [utils.py:54:is_zero_supported_optimizer] Checking ZeRO support for optimizer=FusedAdam type=<class 'deepspeed.ops.adam.fused_adam.FusedAdam'>
[2025-10-22 16:11:08,602] [INFO] [logging.py:96:log_dist] [Rank 0] Creating torch.bfloat16 ZeRO stage 2 optimizer
[2025-10-22 16:11:08,602] [INFO] [stage_1_and_2.py:133:__init__] Reduce bucket size 500000000
[2025-10-22 16:11:08,602] [INFO] [stage_1_and_2.py:134:__init__] Allgather bucket size 500000000
[2025-10-22 16:11:08,602] [INFO] [stage_1_and_2.py:135:__init__] CPU Offload: False
[2025-10-22 16:11:08,602] [INFO] [stage_1_and_2.py:136:__init__] Round robin gradient partitioning: False
Rank: 0 partition count [2] and sizes[(259707438, False)] 
Rank: 1 partition count [2] and sizes[(259707438, False)] 
(train) >> AFTER DEEPSPEED
>> (train) Auto-resume from:  ./runs/plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025/plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025_bidirbio_2048_maxlen512_epochs25_bidir_bio_feedback_loop_train_prompt_enc_srates_9_7_7_1_ckpt_model
>> (train) resume exists:  False
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000246869.jpg', tensor([[[1.1529, 1.1529, 1.1529,  ..., 1.9749, 1.9749, 1.9749],
         [1.1529, 1.1529, 1.1700,  ..., 1.9749, 1.9749, 1.9749],
         [1.1700, 1.1700, 1.1872,  ..., 1.9749, 1.9749, 1.9749],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.4482, 1.4657, 1.4832,  ..., 2.1660, 2.1660, 2.1660],
         [1.4657, 1.4832, 1.5007,  ..., 2.1660, 2.1660, 2.1660],
         [1.5007, 1.5007, 1.5182,  ..., 2.1660, 2.1660, 2.1660],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.8905, 1.8905, 1.9080,  ..., 2.4657, 2.4657, 2.4657],
         [1.9080, 1.9080, 1.9254,  ..., 2.4657, 2.4657, 2.4657],
         [1.9254, 1.9254, 1.9428,  ..., 2.4657, 2.4657, 2.4657],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[1.3610, 1.3610, 1.3902,  ..., 1.4194, 1.4924, 1.5508],
         [1.4048, 1.3756, 1.3610,  ..., 1.4340, 1.4924, 1.5216],
         [1.4194, 1.3902, 1.3902,  ..., 1.4632, 1.5070, 1.5070],
         ...,
         [0.1055, 0.0909, 0.0471,  ..., 0.3537, 0.3537, 0.3537],
         [0.1639, 0.1055, 0.1639,  ..., 0.3683, 0.3829, 0.3975],
         [0.2077, 0.0617, 0.2223,  ..., 0.3245, 0.3391, 0.3391]],

        [[1.6247, 1.6397, 1.6397,  ..., 1.6547, 1.6997, 1.7447],
         [1.6397, 1.6397, 1.6247,  ..., 1.6697, 1.6997, 1.7147],
         [1.6397, 1.6397, 1.6397,  ..., 1.6847, 1.6997, 1.7147],
         ...,
         [0.9343, 0.3490, 0.2740,  ..., 0.8442, 0.8593, 0.8593],
         [1.2044, 0.5291, 0.3340,  ..., 0.7992, 0.7842, 0.7992],
         [1.3095, 0.8292, 0.3640,  ..., 0.8442, 0.8292, 0.8292]],

        [[1.8473, 1.8615, 1.8757,  ..., 1.8615, 1.8899, 1.9184],
         [1.8757, 1.8615, 1.8473,  ..., 1.8757, 1.8899, 1.9042],
         [1.8757, 1.8757, 1.8615,  ..., 1.9042, 1.8899, 1.8899],
         ...,
         [1.5344, 0.6528, 0.5675,  ..., 1.3922, 1.3780, 1.3922],
         [1.8757, 0.9941, 0.5390,  ..., 1.3922, 1.4065, 1.4065],
         [1.9610, 1.3922, 0.7239,  ..., 1.4207, 1.4349, 1.4349]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the the top of a bun in this picture. ASSISTANT: Displayed here is the the top of a bun.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is a sauce covered bun lies next to the rest of a saucy meat sandwich in this image? Please elaborate your answer and explain why. ASSISTANT: I've marked the a sauce covered bun lies next to the rest of a saucy meat sandwich for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the meat on a bun in this image. ASSISTANT: The region corresponding to the meat on a bun is shown.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the a pile of meat with barbecue sauce in this image. ASSISTANT: Result: a pile of meat with barbecue sauce.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nPlease point out the the top of a bun in this picture.', '<image>\nWhat is a sauce covered bun lies next to the rest of a saucy meat sandwich in this image? Please elaborate your answer and explain why.', '<image>\nPlease segment the meat on a bun in this image.', '<image>\nPlease segment the a pile of meat with barbecue sauce in this image.'], ['the top of a bun', 'a sauce covered bun lies next to the rest of a saucy meat sandwich', 'meat on a bun', 'a pile of meat with barbecue sauce'], [{'the top of a bun': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'a sauce covered bun lies next to the rest of a saucy meat sandwich': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'meat on a bun': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'a pile of meat with barbecue sauce': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/p8Rb8E2RXD99WfDJiG1p6A.jpg', tensor([[[ 1.5810,  2.1119,  2.2147,  ..., -1.4500, -1.4672, -1.4672],
         [ 2.0777,  2.1804,  2.2147,  ..., -1.5014, -1.4672, -1.4843],
         [ 1.8550,  2.1290,  2.1975,  ..., -1.4843, -1.4672, -1.4843],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.8508,  2.3761,  2.4111,  ..., -1.0378, -1.1078, -1.1078],
         [ 2.2885,  2.4111,  2.4111,  ..., -1.1254, -1.1254, -1.1604],
         [ 1.9909,  2.2710,  2.3585,  ..., -1.1604, -1.1604, -1.2129],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.1520,  2.5703,  2.5703,  ..., -0.6367, -0.7238, -0.7413],
         [ 2.5354,  2.5877,  2.5529,  ..., -0.7064, -0.7238, -0.7936],
         [ 2.2566,  2.5180,  2.5877,  ..., -0.7238, -0.7587, -0.8284],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.9303,  1.9303,  1.9303,  ...,  1.9303,  1.9303,  1.9303],
         [ 1.9303,  1.9303,  1.9303,  ...,  1.9303,  1.9303,  1.9303],
         [ 1.9303,  1.9303,  1.9303,  ...,  1.9303,  1.9303,  1.9303],
         ...,
         [ 0.1201,  0.1347,  0.1347,  ...,  0.0909,  0.0909,  0.1493],
         [ 0.1055,  0.1201,  0.1201,  ...,  0.1639,  0.1639,  0.1931],
         [ 0.0617,  0.0617,  0.0909,  ..., -0.1718, -0.1426, -0.1280]],

        [[ 2.0749,  2.0749,  2.0749,  ...,  2.0749,  2.0749,  2.0749],
         [ 2.0749,  2.0749,  2.0749,  ...,  2.0749,  2.0749,  2.0749],
         [ 2.0749,  2.0749,  2.0749,  ...,  2.0749,  2.0749,  2.0749],
         ...,
         [ 0.8142,  0.8442,  0.8442,  ...,  0.7392,  0.7392,  0.7992],
         [ 0.7992,  0.8292,  0.8142,  ...,  0.8292,  0.8142,  0.8292],
         [ 0.7842,  0.7842,  0.7842,  ...,  0.4691,  0.4841,  0.4991]],

        [[ 2.1459,  2.1459,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         [ 2.1459,  2.1459,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         [ 2.1459,  2.1459,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         ...,
         [ 1.3069,  1.2927,  1.2927,  ...,  1.1505,  1.1505,  1.2074],
         [ 1.2643,  1.2785,  1.2643,  ...,  1.2358,  1.2216,  1.2643],
         [ 1.2216,  1.2358,  1.2358,  ...,  0.9514,  0.9656,  0.9799]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the vegetation in this image. ASSISTANT: Sure — this shows the vegetation.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is sky in this image? Please give some explanation. ASSISTANT: The sky is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the motorcycle in this picture. ASSISTANT: Displayed here is the motorcycle.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the ego vehicle from the image below? ASSISTANT: The ego vehicle portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is street light in this image? Please give some explanation. ASSISTANT: Below you can see the street light.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[ 61,  61,  61,  ..., 118, 118, 118],
        [ 61,  61,  61,  ..., 118, 118, 118],
        [ 61,  61,  61,  ..., 118, 118, 118],
        ...,
        [120, 120, 120,  ..., 120, 120, 120],
        [120, 120, 120,  ..., 120, 120, 120],
        [120, 120, 120,  ..., 120, 120, 120]]), (576, 1024), ['<image>\nPlease segment the vegetation in this image.', '<image>\nWhat is sky in this image? Please give some explanation.', '<image>\nPlease point out the motorcycle in this picture.', '<image>\nWould you please extract the ego vehicle from the image below?', '<image>\nWhat is street light in this image? Please give some explanation.'], ['vegetation', 'sky', 'motorcycle', 'ego vehicle', 'street light'], [{'vegetation': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'sky': tensor([[ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'motorcycle': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'ego vehicle': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True]])}, {'street light': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000065218.jpg', tensor([[[0.9132, 0.9646, 1.0331,  ..., 0.9303, 0.9646, 0.9817],
         [0.8789, 0.9303, 0.9988,  ..., 0.9303, 0.9646, 0.9646],
         [0.8276, 0.8789, 0.9474,  ..., 0.9303, 0.9474, 0.9474],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.3957, 1.4307, 1.4657,  ..., 1.3431, 1.3782, 1.3957],
         [1.3606, 1.3957, 1.4482,  ..., 1.3431, 1.3782, 1.3782],
         [1.3081, 1.3606, 1.4307,  ..., 1.3431, 1.3606, 1.3606],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.5420, 1.5942, 1.6640,  ..., 1.6465, 1.6814, 1.6988],
         [1.4897, 1.5420, 1.6291,  ..., 1.6465, 1.6814, 1.6814],
         [1.4200, 1.4897, 1.5768,  ..., 1.6465, 1.6640, 1.6640],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 0.3537,  0.3975,  0.4267,  ..., -1.0477, -1.0185, -1.0185],
         [ 0.3099,  0.3245,  0.3683,  ..., -1.0623, -1.0331, -1.0331],
         [ 0.2953,  0.2953,  0.3391,  ..., -1.0915, -1.0477, -1.0769],
         ...,
         [ 1.5946,  1.5946,  1.6092,  ...,  1.2004,  1.2004,  1.2150],
         [ 1.6092,  1.6092,  1.5946,  ...,  1.2442,  1.2442,  1.2588],
         [ 1.6530,  1.6676,  1.6676,  ...,  1.2880,  1.3026,  1.3172]],

        [[ 0.6792,  0.7242,  0.7542,  ..., -0.5365, -0.4614, -0.4164],
         [ 0.6341,  0.6491,  0.6942,  ..., -0.5365, -0.4914, -0.4764],
         [ 0.6191,  0.6191,  0.6642,  ..., -0.5215, -0.4764, -0.5215],
         ...,
         [ 1.7447,  1.7447,  1.7447,  ...,  1.3545,  1.3395,  1.3545],
         [ 1.7597,  1.7597,  1.7447,  ...,  1.3995,  1.3845,  1.3995],
         [ 1.8047,  1.8198,  1.8198,  ...,  1.4145,  1.4446,  1.4596]],

        [[ 0.7808,  0.8234,  0.8377,  ..., -1.2954, -1.3380, -1.4091],
         [ 0.7097,  0.7523,  0.7950,  ..., -1.3522, -1.3665, -1.4091],
         [ 0.7097,  0.6955,  0.7666,  ..., -1.4518, -1.4376, -1.4518],
         ...,
         [ 1.8615,  1.8615,  1.8615,  ...,  1.3496,  1.3638,  1.3780],
         [ 1.8899,  1.8899,  1.8757,  ...,  1.4065,  1.4207,  1.4207],
         [ 1.9468,  1.9610,  1.9610,  ...,  1.4776,  1.4776,  1.4918]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the dining table shown in this image. ASSISTANT: You can see the dining table in this frame.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False,  True,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255],
        [255, 255,  66,  ..., 255, 255, 255],
        ...,
        [255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255]]), (683, 1024), ['<image>\nDisplay a segmentation mask for the dining table shown in this image.'], ['dining table'], [{'dining table': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False,  True,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000375129.jpg', tensor([[[ 1.9235,  1.9235,  1.9235,  ..., -1.9980, -2.0152, -2.0323],
         [ 1.9235,  1.9235,  1.9235,  ..., -1.9980, -2.0152, -2.0323],
         [ 1.9235,  1.9235,  1.9235,  ..., -1.9809, -2.0152, -2.0323],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.0959,  2.0959,  2.0959,  ..., -1.6856, -1.7031, -1.7206],
         [ 2.0959,  2.0959,  2.0959,  ..., -1.6856, -1.7031, -1.7206],
         [ 2.0959,  2.0959,  2.0959,  ..., -1.6681, -1.7031, -1.7206],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.2740,  2.2740,  2.2740,  ..., -1.3513, -1.3687, -1.3861],
         [ 2.2740,  2.2740,  2.2740,  ..., -1.3513, -1.3687, -1.3861],
         [ 2.2740,  2.2740,  2.2740,  ..., -1.3339, -1.3687, -1.3861],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[1.8281, 1.8135, 1.7990,  ..., 1.7260, 1.7406, 1.7406],
         [1.8281, 1.8135, 1.7990,  ..., 1.7260, 1.7406, 1.7406],
         [1.8135, 1.8135, 1.8135,  ..., 1.7406, 1.7406, 1.7260],
         ...,
         [1.2296, 0.8792, 1.5216,  ..., 1.4486, 1.4486, 1.4486],
         [0.9960, 0.8355, 1.0690,  ..., 1.4778, 1.4778, 1.4632],
         [1.0544, 1.1712, 0.9522,  ..., 1.3610, 1.4194, 1.4924]],

        [[2.0299, 2.0149, 1.9998,  ..., 1.9248, 1.9398, 1.9398],
         [2.0299, 2.0149, 1.9998,  ..., 1.9248, 1.9398, 1.9398],
         [2.0149, 2.0149, 2.0149,  ..., 1.9398, 1.9398, 1.9248],
         ...,
         [1.0694, 0.6491, 1.5496,  ..., 1.5646, 1.5796, 1.5646],
         [0.8593, 0.4991, 0.9193,  ..., 1.5946, 1.6096, 1.5946],
         [0.9643, 0.9193, 0.6942,  ..., 1.4746, 1.5496, 1.6397]],

        [[2.1175, 2.1032, 2.0890,  ..., 1.9895, 2.0037, 2.0037],
         [2.1175, 2.1032, 2.0890,  ..., 1.9895, 2.0037, 2.0037],
         [2.1032, 2.1032, 2.1032,  ..., 2.0037, 2.0037, 1.9895],
         ...,
         [1.1078, 0.7239, 1.5629,  ..., 1.7335, 1.7335, 1.7335],
         [0.8803, 0.5817, 0.9514,  ..., 1.7477, 1.7620, 1.7477],
         [0.9656, 0.9656, 0.7239,  ..., 1.6340, 1.7193, 1.7904]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the old man sitting between two other old men wearing a coat in this picture? ASSISTANT: Here is the old man sitting between two other old men wearing a coat you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the the man on the bench to the left in the right hand picture shown in this image. ASSISTANT: The region corresponding to the the man on the bench to the left in the right hand picture is shown.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the brown wooden bench with men sitting on it from the image below? ASSISTANT: Here is the brown wooden bench with men sitting on it you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the bench with three men on it in this image. ASSISTANT: The region corresponding to the bench with three men on it is shown.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is man wearing blue sweater in this image? Please give some explanation. ASSISTANT: You can see the man wearing blue sweater in this frame.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (683, 1024), ['<image>\nCould you isolate the old man sitting between two other old men wearing a coat in this picture?', '<image>\nDisplay a segmentation mask for the the man on the bench to the left in the right hand picture shown in this image.', '<image>\nWould you please extract the brown wooden bench with men sitting on it from the image below?', '<image>\nPlease segment the bench with three men on it in this image.', '<image>\nWhat is man wearing blue sweater in this image? Please give some explanation.'], ['old man sitting between two other old men wearing a coat', 'the man on the bench to the left in the right hand picture', 'brown wooden bench with men sitting on it', 'bench with three men on it', 'man wearing blue sweater'], [{'old man sitting between two other old men wearing a coat': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the man on the bench to the left in the right hand picture': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'brown wooden bench with men sitting on it': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'bench with three men on it': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'man wearing blue sweater': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000437282.jpg', tensor([[[-1.2103, -1.1075, -0.9705,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2103, -1.1075, -0.9877,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2103, -1.1247, -1.0048,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.0458, -0.0458, -0.0458,  ...,  0.0000,  0.0000,  0.0000],
         [-0.0801, -0.0801, -0.0801,  ...,  0.0000,  0.0000,  0.0000],
         [-0.0972, -0.0972, -0.0972,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.9503, -0.8452, -0.7052,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9503, -0.8452, -0.7227,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9503, -0.8627, -0.7402,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.4776, -0.4776, -0.4776,  ...,  0.0000,  0.0000,  0.0000],
         [-0.5126, -0.5126, -0.5126,  ...,  0.0000,  0.0000,  0.0000],
         [-0.5301, -0.5301, -0.5301,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.5670, -0.4624, -0.3230,  ...,  0.0000,  0.0000,  0.0000],
         [-0.5670, -0.4624, -0.3404,  ...,  0.0000,  0.0000,  0.0000],
         [-0.5670, -0.4798, -0.3578,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.9330, -0.9330, -0.9330,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9504, -0.9504, -0.9504,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9678, -0.9678, -0.9678,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.2077,  0.2661, -0.0842,  ...,  0.6603,  0.6895,  0.6895],
         [ 1.3172,  1.1712,  0.9522,  ...,  0.5289,  0.5435,  0.5143],
         [ 1.2150,  1.5800,  1.4632,  ...,  0.2807,  0.2953,  0.2661],
         ...,
         [ 0.4121,  0.3975,  0.3975,  ...,  0.5143,  0.6019,  0.6019],
         [ 0.4121,  0.4121,  0.3975,  ...,  0.4121,  0.4705,  0.5581],
         [ 0.3975,  0.3975,  0.3829,  ...,  0.4705,  0.4267,  0.4705]],

        [[ 0.5591,  0.5141,  0.1089,  ...,  0.7992,  0.8292,  0.8292],
         [ 1.4896,  1.4746,  1.2495,  ...,  0.6942,  0.7242,  0.7392],
         [ 1.4746,  1.7747,  1.7297,  ...,  0.4090,  0.4390,  0.4240],
         ...,
         [ 0.5291,  0.5441,  0.5591,  ...,  0.6792,  0.8292,  0.8893],
         [ 0.5141,  0.5441,  0.5591,  ...,  0.5891,  0.6642,  0.8292],
         [ 0.4991,  0.5291,  0.5441,  ...,  0.5891,  0.5741,  0.6792]],

        [[ 0.8661,  0.9514,  0.6101,  ...,  0.9941,  1.0083,  1.0367],
         [ 1.8615,  1.8473,  1.6624,  ...,  0.8803,  0.9088,  0.9088],
         [ 1.5913,  1.8331,  1.7477,  ...,  0.6244,  0.6528,  0.6386],
         ...,
         [ 0.7381,  0.7381,  0.7523,  ...,  0.9656,  1.1363,  1.1647],
         [ 0.7239,  0.7381,  0.7523,  ...,  0.8092,  0.9088,  1.0794],
         [ 0.7239,  0.7381,  0.7381,  ...,  0.8234,  0.8234,  0.9372]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the key elements in this picture? ASSISTANT: The image shows a young child climbing on the shelves inside a small refrigerator, as if engaged in a playful adventure. The refrigerator door is open, revealing various food items stored within. There is a multitude of bottles placed on different shelves, an orange in one of the compartments, and a few carrots scattered inside.\n\nFurthermore, several bowls can be seen positioned across the storage compartments, containing different foods. The child appears to be standing on the bottom shelf, likely enjoying the unusual experience of exploring the inside of the fridge.</s>"], tensor([], size=(0, 640, 426)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 682), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the key elements in this picture? ASSISTANT: The image shows a young child climbing on the shelves inside a small refrigerator, as if engaged in a playful adventure. The refrigerator door is open, revealing various food items stored within. There is a multitude of bottles placed on different shelves, an orange in one of the compartments, and a few carrots scattered inside.\n\nFurthermore, several bowls can be seen positioned across the storage compartments, containing different foods. The child appears to be standing on the bottom shelf, likely enjoying the unusual experience of exploring the inside of the fridge.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the key elements in this picture? ASSISTANT: The image shows a young child climbing on the shelves inside a small refrigerator, as if engaged in a playful adventure. The refrigerator door is open, revealing various food items stored within. There is a multitude of bottles placed on different shelves, an orange in one of the compartments, and a few carrots scattered inside.\n\nFurthermore, several bowls can be seen positioned across the storage compartments, containing different foods. The child appears to be standing on the bottom shelf, likely enjoying the unusual experience of exploring the inside of the fridge.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000282039.jpg', tensor([[[ 0.9817,  0.9817,  0.9817,  ...,  1.6838,  1.6667,  1.6667],
         [ 0.9817,  0.9817,  0.9817,  ...,  1.6838,  1.6667,  1.6667],
         [ 0.9988,  0.9988,  0.9988,  ...,  1.6667,  1.6667,  1.6667],
         ...,
         [-1.4672, -1.4672, -1.4672,  ...,  0.6734,  0.6563,  0.6563],
         [-1.5014, -1.5014, -1.5014,  ...,  0.7077,  0.7077,  0.7077],
         [-1.5185, -1.5185, -1.5185,  ...,  0.7419,  0.7419,  0.7419]],

        [[ 1.5357,  1.5357,  1.5357,  ...,  2.0434,  2.0259,  2.0259],
         [ 1.5357,  1.5357,  1.5357,  ...,  2.0434,  2.0259,  2.0259],
         [ 1.5532,  1.5532,  1.5532,  ...,  2.0259,  2.0259,  2.0259],
         ...,
         [-1.7906, -1.7906, -1.7906,  ...,  0.5553,  0.5378,  0.5378],
         [-1.8256, -1.8256, -1.8256,  ...,  0.5903,  0.5903,  0.5903],
         [-1.8431, -1.8431, -1.8431,  ...,  0.6254,  0.6254,  0.6254]],

        [[ 1.4374,  1.4374,  1.4374,  ...,  1.6640,  1.6465,  1.6465],
         [ 1.4374,  1.4374,  1.4374,  ...,  1.6640,  1.6465,  1.6465],
         [ 1.4548,  1.4548,  1.4548,  ...,  1.6465,  1.6465,  1.6465],
         ...,
         [-1.1770, -1.1770, -1.1770,  ...,  0.4439,  0.4265,  0.4265],
         [-1.2119, -1.2119, -1.2119,  ...,  0.4788,  0.4788,  0.4788],
         [-1.2293, -1.2293, -1.2293,  ...,  0.5136,  0.5136,  0.5136]]]), tensor([[[ 0.8792,  0.9230,  0.9522,  ...,  1.4340,  1.4486,  1.4486],
         [ 0.9084,  0.9230,  0.9376,  ...,  1.4486,  1.4632,  1.4486],
         [ 0.9376,  0.9230,  0.9376,  ...,  1.4486,  1.4632,  1.4632],
         ...,
         [-0.7558, -0.7412, -0.7412,  ...,  0.3391,  0.3245,  0.2807],
         [-1.0331, -1.0185, -1.0185,  ...,  0.4121,  0.4413,  0.3975],
         [-1.2229, -1.2229, -1.2229,  ...,  0.5143,  0.5727,  0.5727]],

        [[ 1.3395,  1.3845,  1.4145,  ...,  1.7147,  1.7447,  1.7447],
         [ 1.3695,  1.3845,  1.3995,  ...,  1.7297,  1.7597,  1.7447],
         [ 1.3995,  1.3845,  1.3995,  ...,  1.7297,  1.7597,  1.7597],
         ...,
         [-0.9267, -0.9117, -0.9117,  ...,  0.1689,  0.1389,  0.0638],
         [-1.3019, -1.2718, -1.2869,  ...,  0.2740,  0.2890,  0.2439],
         [-1.5270, -1.5270, -1.5270,  ...,  0.3940,  0.4540,  0.4691]],

        [[ 1.1932,  1.2358,  1.2643,  ...,  1.3211,  1.3496,  1.3496],
         [ 1.2216,  1.2358,  1.2500,  ...,  1.3354,  1.3638,  1.3496],
         [ 1.2500,  1.2358,  1.2500,  ...,  1.3354,  1.3638,  1.3638],
         ...,
         [-0.4706, -0.4422, -0.4422,  ...,  0.0555,  0.0698,  0.0413],
         [-0.7550, -0.7123, -0.7123,  ...,  0.1835,  0.1977,  0.1693],
         [-0.9541, -0.9114, -0.8972,  ...,  0.3115,  0.3399,  0.3399]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat might the dog be feeling or thinking as it looks through the blinds? ASSISTANT: Although it's impossible to know the exact thoughts and feelings of the dog, we can attempt to infer based on its actions. As the white Westie dog is looking through the blinds out the window, it might be feeling curious, vigilant, or excited about the activities happening in the outside world. Dogs often look outside to watch for approaching people or animals, to observe their surroundings, or to satisfy their natural curiosity about the world around them. The dog could also be waiting for someone it knows, such as its owner, and it might feel anticipation or impatience. However, any assumption made about the dog's emotions or thoughts is speculative and open to interpretation.</s>"], tensor([], size=(0, 612, 612)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat might the dog be feeling or thinking as it looks through the blinds? ASSISTANT: Although it's impossible to know the exact thoughts and feelings of the dog, we can attempt to infer based on its actions. As the white Westie dog is looking through the blinds out the window, it might be feeling curious, vigilant, or excited about the activities happening in the outside world. Dogs often look outside to watch for approaching people or animals, to observe their surroundings, or to satisfy their natural curiosity about the world around them. The dog could also be waiting for someone it knows, such as its owner, and it might feel anticipation or impatience. However, any assumption made about the dog's emotions or thoughts is speculative and open to interpretation.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat might the dog be feeling or thinking as it looks through the blinds? ASSISTANT: Although it's impossible to know the exact thoughts and feelings of the dog, we can attempt to infer based on its actions. As the white Westie dog is looking through the blinds out the window, it might be feeling curious, vigilant, or excited about the activities happening in the outside world. Dogs often look outside to watch for approaching people or animals, to observe their surroundings, or to satisfy their natural curiosity about the world around them. The dog could also be waiting for someone it knows, such as its owner, and it might feel anticipation or impatience. However, any assumption made about the dog's emotions or thoughts is speculative and open to interpretation.</s>"], [{}], False, 'vqa')]
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000330429.jpg', tensor([[[ 1.5639,  1.5468,  1.5297,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.5297,  1.5468,  1.5982,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.4954,  1.5639,  1.6667,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 1.3242,  1.3070,  1.2899,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.4612,  1.4269,  1.3584,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.5810,  1.5125,  1.4098,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.4503,  0.3803,  0.2927,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.4328,  0.4153,  0.3803,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.4153,  0.4503,  0.5028,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 0.2927,  0.2752,  0.2577,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.3978,  0.3627,  0.2927,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.4853,  0.4153,  0.3277,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.7064, -0.4798, -0.2010,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9156, -0.6715, -0.3753,  ...,  0.0000,  0.0000,  0.0000],
         [-1.1421, -0.8981, -0.5670,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.1596, -1.1596, -1.1944,  ...,  0.0000,  0.0000,  0.0000],
         [-0.8284, -0.8981, -1.0027,  ...,  0.0000,  0.0000,  0.0000],
         [-0.5495, -0.6715, -0.8284,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.4632,  1.5362,  1.5070,  ...,  1.5946,  1.5946,  1.5508],
         [ 1.4778,  1.4924,  1.5654,  ...,  1.6092,  1.6238,  1.5946],
         [ 1.5216,  1.5216,  1.5216,  ...,  1.6384,  1.6238,  1.6238],
         ...,
         [ 1.4194,  1.4048,  1.4194,  ...,  1.4924,  1.5070,  1.5070],
         [ 1.3756,  1.3610,  1.4048,  ...,  1.4924,  1.5362,  1.5070],
         [ 1.4048,  1.3756,  1.3902,  ...,  1.5362,  1.5654,  1.5070]],

        [[ 0.5591,  0.4991,  0.4390,  ...,  0.6341,  0.6491,  0.6642],
         [ 0.5291,  0.4090,  0.4691,  ...,  0.6341,  0.6491,  0.6642],
         [ 0.4841,  0.4390,  0.4240,  ...,  0.6792,  0.6341,  0.6341],
         ...,
         [ 0.3640,  0.3940,  0.3040,  ...,  0.6191,  0.5891,  0.5141],
         [ 0.4390,  0.3790,  0.4390,  ...,  0.5741,  0.5441,  0.5591],
         [ 0.3940,  0.3190,  0.3790,  ...,  0.5591,  0.4991,  0.5741]],

        [[-0.8545, -0.5275, -0.6412,  ..., -0.2289, -0.1435, -0.0298],
         [-0.6839, -0.3995, -0.5275,  ..., -0.2289, -0.2004, -0.0440],
         [-0.2004, -0.2573, -0.3995,  ..., -0.2431, -0.3426, -0.1578],
         ...,
         [-0.4564, -0.8688, -0.6555,  ..., -0.5559, -0.2715, -0.2146],
         [-0.3853, -0.6981, -0.8261,  ..., -0.6412, -0.3284, -0.3711],
         [-0.3426, -0.5417, -0.5559,  ..., -0.3284, -0.0867, -0.3284]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat improvements could be made to the bathroom? ASSISTANT: Looking at the image, the bathroom appears quite simple and rundown. To enhance its appearance and functionality, several improvements could be made. First, the bathroom has no visible floor, so installing a proper flooring like tiles or laminate would significantly improve the space. Adding storage options, such as shelves or cabinets, would help organize toiletries and provide a clutter-free environment. Additionally, replacing old or damaged fixtures, like faucets, showerheads, or light fixtures, with modern designs would add an updated touch. The inclusion of a mirror or upgrading the existing one, if there is one, could create a more polished, cohesive space. Finally, incorporating various decorative elements like colorful towels, shower curtains, or wall art could enhance the overall aesthetic and make the bathroom more inviting.</s>"], tensor([], size=(0, 640, 480)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 768), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat improvements could be made to the bathroom? ASSISTANT: Looking at the image, the bathroom appears quite simple and rundown. To enhance its appearance and functionality, several improvements could be made. First, the bathroom has no visible floor, so installing a proper flooring like tiles or laminate would significantly improve the space. Adding storage options, such as shelves or cabinets, would help organize toiletries and provide a clutter-free environment. Additionally, replacing old or damaged fixtures, like faucets, showerheads, or light fixtures, with modern designs would add an updated touch. The inclusion of a mirror or upgrading the existing one, if there is one, could create a more polished, cohesive space. Finally, incorporating various decorative elements like colorful towels, shower curtains, or wall art could enhance the overall aesthetic and make the bathroom more inviting.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat improvements could be made to the bathroom? ASSISTANT: Looking at the image, the bathroom appears quite simple and rundown. To enhance its appearance and functionality, several improvements could be made. First, the bathroom has no visible floor, so installing a proper flooring like tiles or laminate would significantly improve the space. Adding storage options, such as shelves or cabinets, would help organize toiletries and provide a clutter-free environment. Additionally, replacing old or damaged fixtures, like faucets, showerheads, or light fixtures, with modern designs would add an updated touch. The inclusion of a mirror or upgrading the existing one, if there is one, could create a more polished, cohesive space. Finally, incorporating various decorative elements like colorful towels, shower curtains, or wall art could enhance the overall aesthetic and make the bathroom more inviting.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/HPbpY74NDMME0F4X4OdgVw.jpg', tensor([[[ 2.2489,  2.2489,  2.2489,  ..., -0.7822, -0.7993, -0.7650],
         [ 2.2489,  2.2489,  2.2489,  ..., -0.8335, -0.7822, -0.7822],
         [ 2.2489,  2.2489,  2.2489,  ..., -0.7479, -0.7137, -0.7308],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.4286,  2.4286,  2.4286,  ..., -0.7752, -0.7927, -0.7577],
         [ 2.4286,  2.4286,  2.4286,  ..., -0.8277, -0.7752, -0.7752],
         [ 2.4286,  2.4286,  2.4286,  ..., -0.7402, -0.7052, -0.7227],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.6400,  2.6400,  2.6400,  ..., -0.3055, -0.3230, -0.2881],
         [ 2.6400,  2.6400,  2.6400,  ..., -0.3578, -0.3055, -0.3055],
         [ 2.6400,  2.6400,  2.6400,  ..., -0.2707, -0.2358, -0.2532],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.9303,  1.9303,  1.9303,  ...,  0.3537,  0.4121,  0.6165],
         [ 1.9303,  1.9303,  1.9303,  ...,  0.3975,  0.4413,  0.4705],
         [ 1.9303,  1.9303,  1.9303,  ...,  0.5143,  0.6165,  0.3683],
         ...,
         [ 0.2661,  0.2223,  0.0909,  ..., -0.5368, -0.5514, -0.6098],
         [-0.0259, -0.1280,  0.0471,  ..., -0.6390, -0.6828, -1.0769],
         [-0.1572, -0.0113,  0.1785,  ..., -0.6098, -1.0623, -1.0185]],

        [[ 2.0749,  2.0749,  2.0749,  ...,  0.3940,  0.3940,  0.7692],
         [ 2.0749,  2.0749,  2.0749,  ...,  0.4540,  0.5891,  0.6642],
         [ 2.0749,  2.0749,  2.0749,  ...,  0.5141,  0.8292,  0.6191],
         ...,
         [ 0.4540,  0.3790,  0.2139,  ..., -0.4464, -0.4614, -0.5365],
         [ 0.1239,  0.0038,  0.1539,  ..., -0.5515, -0.5965, -1.0467],
         [-0.0112,  0.1389,  0.3190,  ..., -0.5215, -1.0017, -0.9867]],

        [[ 2.1459,  2.1459,  2.1459,  ...,  0.6386,  0.6812,  0.9799],
         [ 2.1459,  2.1459,  2.1459,  ...,  0.7097,  0.8092,  0.8234],
         [ 2.1459,  2.1459,  2.1459,  ...,  0.8519,  1.0652,  0.7381],
         ...,
         [ 0.3684,  0.3257,  0.2120,  ..., -0.3426, -0.3284, -0.3711],
         [ 0.0982, -0.0156,  0.1409,  ..., -0.4422, -0.4422, -0.8545],
         [-0.0440,  0.0982,  0.2688,  ..., -0.3853, -0.7550, -0.7834]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the signage - advertisement shown in this image. ASSISTANT: This is the selected signage - advertisement.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the car in this image. ASSISTANT: Take a look at the car here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the manhole in this picture? ASSISTANT: Take a look at the manhole here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the lane marking - straight line in this image? ASSISTANT: The lane marking - straight line portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the sky from the image below? ASSISTANT: The region corresponding to the sky is shown.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[61, 61, 61,  ..., 64, 64, 64],
        [61, 61, 61,  ..., 64, 64, 64],
        [61, 61, 61,  ..., 64, 64, 64],
        ...,
        [24, 24, 24,  ..., 24, 24, 24],
        [24, 24, 24,  ..., 24, 24, 24],
        [24, 24, 24,  ..., 24, 24, 24]]), (765, 1024), ['<image>\nDisplay a segmentation mask for the signage - advertisement shown in this image.', '<image>\nPlease highlight the car in this image.', '<image>\nCould you identify and segment out the manhole in this picture?', '<image>\nCan you segment the lane marking - straight line in this image?', '<image>\nWould you please extract the sky from the image below?'], ['signage - advertisement', 'car', 'manhole', 'lane marking - straight line', 'sky'], [{'signage - advertisement': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'car': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'manhole': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'lane marking - straight line': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'sky': tensor([[ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000228893.jpg', tensor([[[-1.1247, -1.1418, -1.1589,  ..., -0.7650, -0.9363, -1.0562],
         [-1.2103, -1.1932, -1.1589,  ..., -0.7993, -0.7993, -0.7993],
         [-1.3302, -1.2617, -1.1760,  ..., -0.8335, -0.6281, -0.4739],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.1275, -0.1800, -0.2500,  ..., -0.3200, -0.4776, -0.6001],
         [-0.2500, -0.3025, -0.3901,  ..., -0.4076, -0.4601, -0.5126],
         [-0.4251, -0.4951, -0.5651,  ..., -0.5126, -0.4426, -0.3901],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.5081, -1.4733, -1.4210,  ..., -0.4798, -0.6890, -0.8633],
         [-1.4733, -1.4733, -1.4559,  ..., -0.6367, -0.6367, -0.6541],
         [-1.4384, -1.4733, -1.4907,  ..., -0.8633, -0.5844, -0.3927],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.3762, -0.3032, -0.3324,  ..., -0.3908, -0.0842, -0.4346],
         [-0.2740, -0.2740, -0.4054,  ..., -0.4200, -0.2594, -0.4346],
         [-0.2302, -0.1426, -0.2448,  ..., -0.3616, -0.3616, -0.5514],
         ...,
         [-0.4346, -0.5660, -0.7266,  ..., -0.6828, -0.6536, -0.7850],
         [-0.4638, -0.5514, -0.6098,  ..., -0.6536, -0.6974, -0.8580],
         [-0.5660, -0.5222, -0.4930,  ..., -0.7704, -0.6682, -0.7266]],

        [[ 0.1689,  0.2589,  0.1689,  ...,  0.0338,  0.2139, -0.0112],
         [ 0.3040,  0.3640,  0.2289,  ...,  0.0939,  0.2289,  0.0488],
         [ 0.5291,  0.5291,  0.3790,  ...,  0.1839,  0.1239,  0.0789],
         ...,
         [ 0.2139,  0.1239, -0.0262,  ...,  0.1839,  0.1689, -0.0262],
         [ 0.1239,  0.0939,  0.0338,  ...,  0.3190,  0.1989, -0.0262],
         [ 0.0338,  0.0789,  0.0488,  ...,  0.1089,  0.1689,  0.0939]],

        [[-0.5986, -0.6555, -0.5133,  ..., -0.4279, -0.1009, -0.5417],
         [-0.3853, -0.5559, -0.5986,  ..., -0.5559, -0.2431, -0.4137],
         [-0.6128, -0.3995, -0.5275,  ..., -0.6697, -0.6412, -0.4706],
         ...,
         [-0.3853, -0.6697, -0.7692,  ..., -0.7550, -0.6412, -0.7550],
         [-0.3853, -0.6270, -0.7266,  ..., -0.6555, -0.5986, -0.7692],
         [-0.4990, -0.5133, -0.5986,  ..., -0.7834, -0.5701, -0.6697]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the zebra in this picture. ASSISTANT: Here's where the zebra appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the bush in this image. ASSISTANT: Here's where the bush appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the grass in this image. ASSISTANT: The grass portion is displayed below.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True]]]), tensor([[ 96,  96,  96,  ...,  96,  96,  96],
        [ 96,  96,  96,  ...,  96,  96,  96],
        [ 96,  96,  96,  ...,  96,  96,  96],
        ...,
        [123, 123, 123,  ..., 123, 123, 123],
        [123, 123, 123,  ..., 123, 123, 123],
        [123, 123, 123,  ..., 123, 123, 123]]), (680, 1024), ['<image>\nPlease point out the zebra in this picture.', '<image>\nPlease highlight the bush in this image.', '<image>\nPlease segment the grass in this image.'], ['zebra', 'bush', 'grass'], [{'zebra': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'bush': tensor([[ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'grass': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/vV867h5ECeIETaRoec3_0w.jpg', tensor([[[0.3652, 0.3138, 0.3652,  ..., 2.2489, 2.2489, 2.2489],
         [0.3309, 0.8789, 1.1187,  ..., 2.2489, 2.2489, 2.2489],
         [0.7419, 1.0159, 0.7933,  ..., 2.2489, 2.2489, 2.2489],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.7654, 0.7129, 0.5028,  ..., 2.4286, 2.4286, 2.4286],
         [0.6954, 1.3081, 1.3256,  ..., 2.4286, 2.4286, 2.4286],
         [1.0630, 1.4657, 1.0805,  ..., 2.4286, 2.4286, 2.4286],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.2457, 1.1062, 0.9145,  ..., 2.6400, 2.6400, 2.6400],
         [1.1237, 1.6814, 1.7163,  ..., 2.6400, 2.6400, 2.6400],
         [1.4722, 1.8208, 1.4374,  ..., 2.6400, 2.6400, 2.6400],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[-0.0259, -0.3762,  0.2661,  ...,  1.9303,  1.9303,  1.9303],
         [-0.4200, -0.3908, -0.0550,  ...,  1.9303,  1.9303,  1.9303],
         [-0.7558, -0.3178, -0.2302,  ...,  1.9303,  1.9303,  1.9303],
         ...,
         [-0.9456, -1.0039, -1.0331,  ..., -0.1572, -0.0988, -0.0550],
         [-0.9893, -0.9893, -0.9893,  ..., -0.0696, -0.1280, -0.1134],
         [-1.0039, -0.9602, -1.0039,  ..., -0.0842, -0.1134, -0.1280]],

        [[ 0.0038, -0.4614,  0.3490,  ...,  2.0749,  2.0749,  2.0749],
         [-0.4164, -0.5065, -0.0712,  ...,  2.0749,  2.0749,  2.0749],
         [-0.7766, -0.3564, -0.1913,  ...,  2.0749,  2.0749,  2.0749],
         ...,
         [-0.9117, -0.9417, -0.9717,  ..., -0.2513, -0.1913, -0.1463],
         [-0.9717, -0.9267, -0.9267,  ..., -0.1613, -0.2063, -0.2063],
         [-0.9717, -0.8967, -0.9417,  ..., -0.1313, -0.1763, -0.2063]],

        [[ 0.1977, -0.3142,  0.5532,  ...,  2.1459,  2.1459,  2.1459],
         [-0.2431, -0.3142,  0.1409,  ...,  2.1459,  2.1459,  2.1459],
         [-0.6270, -0.2715,  0.0840,  ...,  2.1459,  2.1459,  2.1459],
         ...,
         [-0.6412, -0.6839, -0.7123,  ..., -0.1862, -0.1435, -0.1151],
         [-0.6839, -0.6697, -0.6697,  ..., -0.1435, -0.1578, -0.1862],
         [-0.6697, -0.6270, -0.6839,  ..., -0.1293, -0.1293, -0.1862]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the lane marking - dashed line in this picture. ASSISTANT: Below you can see the lane marking - dashed line.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the lane marking - other in this image? ASSISTANT: This view focuses on the lane marking - other.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the pole in this image? ASSISTANT: I've marked the pole for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the ground in this picture? ASSISTANT: Result: ground.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the ego vehicle shown in this image. ASSISTANT: Result: ego vehicle.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False]]]), tensor([[ 64,  64,  64,  ...,  61,  61,  61],
        [ 64,  64,  64,  ...,  61,  61,  61],
        [ 64,  64,  64,  ...,  61,  61,  61],
        ...,
        [120, 120, 120,  ...,  21,  21,  21],
        [120, 120, 120,  ...,  21,  21,  21],
        [120, 120, 120,  ...,  21,  21,  21]]), (768, 1024), ['<image>\nPlease point out the lane marking - dashed line in this picture.', '<image>\nCan you segment the lane marking - other in this image?', '<image>\nCan you segment the pole in this image?', '<image>\nCould you identify and segment out the ground in this picture?', '<image>\nDisplay a segmentation mask for the ego vehicle shown in this image.'], ['lane marking - dashed line', 'lane marking - other', 'pole', 'ground', 'ego vehicle'], [{'lane marking - dashed line': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'lane marking - other': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'pole': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'ground': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'ego vehicle': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000424669.jpg', tensor([[[ 2.1462,  2.1804,  2.2318,  ..., -1.9467, -1.9980, -2.0494],
         [ 2.0605,  2.1119,  2.1804,  ..., -1.9467, -1.9809, -2.0323],
         [ 1.9407,  2.0263,  2.1290,  ..., -1.9295, -1.9638, -1.9980],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.2556,  1.3431,  1.4657,  ..., -1.8256, -1.8782, -1.9307],
         [ 1.1856,  1.2906,  1.4132,  ..., -1.8256, -1.8606, -1.9132],
         [ 1.0980,  1.2031,  1.3606,  ..., -1.8081, -1.8431, -1.8782],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.6531,  0.6531,  0.6705,  ..., -1.6127, -1.6650, -1.7173],
         [ 0.5659,  0.5834,  0.6182,  ..., -1.6127, -1.6476, -1.6999],
         [ 0.4439,  0.4962,  0.5659,  ..., -1.5953, -1.6302, -1.6650],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.4559,  0.3975,  0.4413,  ..., -0.2740, -1.4273, -1.6463],
         [ 0.5143,  0.4267,  0.4267,  ..., -1.6171, -1.6609, -1.6755],
         [ 0.4705,  0.4559,  0.4559,  ..., -1.6609, -1.6171, -1.6463],
         ...,
         [-1.1645, -1.2667, -0.6974,  ..., -1.2521, -1.2521, -1.2083],
         [-0.8580, -1.6025, -1.1353,  ..., -1.1791, -1.2083, -1.2521],
         [-0.5076, -1.6317, -1.3397,  ..., -1.2375, -1.2813, -1.2959]],

        [[ 0.4090,  0.4090,  0.3940,  ..., -0.0712, -1.2869, -1.5720],
         [ 0.4390,  0.4691,  0.4240,  ..., -1.5270, -1.5870, -1.6470],
         [ 0.4390,  0.4841,  0.4390,  ..., -1.6170, -1.5870, -1.6020],
         ...,
         [-1.0617, -1.1818, -0.5665,  ..., -1.1368, -1.1518, -1.1218],
         [-0.6715, -1.4970, -0.9567,  ..., -1.1218, -1.1368, -1.1218],
         [-0.2963, -1.5120, -1.2718,  ..., -1.1368, -1.1668, -1.1668]],

        [[ 0.4110,  0.4110,  0.4395,  ...,  0.4821, -0.7408, -1.2527],
         [ 0.4821,  0.4253,  0.4253,  ..., -1.1816, -1.3096, -1.3807],
         [ 0.4821,  0.4679,  0.4679,  ..., -1.4376, -1.3096, -1.3380],
         ...,
         [-0.8261, -0.7977, -0.2715,  ..., -0.9683, -0.9399, -0.9825],
         [-0.4706, -1.0394, -0.5417,  ..., -0.8830, -0.8830, -1.0110],
         [-0.2004, -1.1674, -0.8688,  ..., -0.9683, -0.8972, -0.9541]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat can you infer about the purpose of this event? ASSISTANT: Based on the image, where people are playing video games on a television screen in a booth with chairs, laptops, and multiple people around, it can be inferred that the purpose of this event is likely to promote, display, or showcase video games, gaming technology, or related products. This could be a trade show or a gaming convention, where attendees can interact with and get hands-on experience with the games and technology on display. The event aims to create an engaging and interactive environment for people to learn about, play, and appreciate the games and technology in the industry.</s>"], tensor([], size=(0, 480, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat can you infer about the purpose of this event? ASSISTANT: Based on the image, where people are playing video games on a television screen in a booth with chairs, laptops, and multiple people around, it can be inferred that the purpose of this event is likely to promote, display, or showcase video games, gaming technology, or related products. This could be a trade show or a gaming convention, where attendees can interact with and get hands-on experience with the games and technology on display. The event aims to create an engaging and interactive environment for people to learn about, play, and appreciate the games and technology in the industry.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat can you infer about the purpose of this event? ASSISTANT: Based on the image, where people are playing video games on a television screen in a booth with chairs, laptops, and multiple people around, it can be inferred that the purpose of this event is likely to promote, display, or showcase video games, gaming technology, or related products. This could be a trade show or a gaming convention, where attendees can interact with and get hands-on experience with the games and technology on display. The event aims to create an engaging and interactive environment for people to learn about, play, and appreciate the games and technology in the industry.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000490610.jpg', tensor([[[-1.0562, -1.0219, -0.9705,  ...,  0.0000,  0.0000,  0.0000],
         [-1.0219, -1.0048, -0.9705,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9877, -0.9877, -0.9877,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.1486, -0.3712, -0.5938,  ...,  0.0000,  0.0000,  0.0000],
         [-0.4226, -0.5082, -0.5767,  ...,  0.0000,  0.0000,  0.0000],
         [-0.6452, -0.6281, -0.5596,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.9678, -0.9153, -0.8627,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9503, -0.9153, -0.8803,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9153, -0.9153, -0.8978,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.3375, -0.5476, -0.7927,  ...,  0.0000,  0.0000,  0.0000],
         [-0.6176, -0.7052, -0.7752,  ...,  0.0000,  0.0000,  0.0000],
         [-0.8452, -0.8277, -0.7577,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.7761, -0.7064, -0.6193,  ...,  0.0000,  0.0000,  0.0000],
         [-0.7238, -0.6715, -0.6193,  ...,  0.0000,  0.0000,  0.0000],
         [-0.6715, -0.6367, -0.6018,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.3578, -0.5844, -0.8110,  ...,  0.0000,  0.0000,  0.0000],
         [-0.6367, -0.7238, -0.7936,  ...,  0.0000,  0.0000,  0.0000],
         [-0.8633, -0.8458, -0.7761,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.4565, -1.4565, -1.4419,  ..., -0.7558, -0.9164, -0.9456],
         [-1.3981, -1.3835, -1.3835,  ..., -0.8288, -0.8288, -0.8872],
         [-1.3835, -1.3981, -1.3835,  ..., -0.8872, -1.0477, -1.1207],
         ...,
         [-0.0550, -0.0988, -0.1864,  ..., -0.2448, -0.2594, -0.2448],
         [-0.0405,  0.0033, -0.0550,  ..., -0.2886, -0.2302, -0.2740],
         [-0.0550,  0.0617, -0.0259,  ..., -0.2594, -0.2886, -0.2594]],

        [[-1.4069, -1.4069, -1.3919,  ..., -0.6565, -0.8366, -0.9117],
         [-1.3619, -1.3319, -1.3319,  ..., -0.7616, -0.7766, -0.7916],
         [-1.3169, -1.3019, -1.3319,  ..., -0.8066, -0.9417, -0.9417],
         ...,
         [-0.3264, -0.4014, -0.4014,  ..., -1.0918, -1.1068, -1.0918],
         [-0.3264, -0.2513, -0.2963,  ..., -1.1668, -1.0918, -1.1218],
         [-0.3864, -0.2813, -0.2663,  ..., -1.1368, -1.1818, -1.1668]],

        [[-1.0394, -1.0394, -1.0252,  ..., -0.5417, -0.5701, -0.5275],
         [-0.9967, -0.9825, -0.9825,  ..., -0.5844, -0.5559, -0.5986],
         [-0.9967, -0.9967, -1.0110,  ..., -0.5559, -0.6412, -0.6697],
         ...,
         [-0.4279, -0.4564, -0.4990,  ..., -0.5986, -0.6128, -0.6270],
         [-0.4137, -0.3568, -0.4279,  ..., -0.6839, -0.6128, -0.6555],
         [-0.4848, -0.3853, -0.3995,  ..., -0.6697, -0.6981, -0.6981]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the towel in this picture. ASSISTANT: Result: towel.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is towel body in this image? Please elaborate your answer and explain why. ASSISTANT: Take a look at the towel body here.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 683), ['<image>\nPlease point out the towel in this picture.', '<image>\nWhat is towel body in this image? Please elaborate your answer and explain why.'], ['towel', 'towel body'], [{'towel': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'towel body': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg')]
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000162452.jpg', tensor([[[-1.7412, -1.7583, -1.8097,  ..., -1.7925, -1.7754, -1.7754],
         [-1.7583, -1.7583, -1.7925,  ..., -1.7925, -1.7754, -1.7754],
         [-1.7925, -1.7754, -1.7583,  ..., -1.8097, -1.7925, -1.7925],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.6506, -1.6681, -1.7206,  ..., -1.7031, -1.6856, -1.6856],
         [-1.6681, -1.6681, -1.7031,  ..., -1.7031, -1.6856, -1.6856],
         [-1.7031, -1.6856, -1.6681,  ..., -1.7206, -1.7031, -1.7031],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.4210, -1.4384, -1.4907,  ..., -1.4733, -1.4559, -1.4559],
         [-1.4384, -1.4384, -1.4733,  ..., -1.4733, -1.4559, -1.4559],
         [-1.4733, -1.4559, -1.4384,  ..., -1.4907, -1.4733, -1.4733],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.5149, -1.3543, -1.2229,  ..., -1.4565, -1.3835, -1.4711],
         [-1.1061, -1.2813, -1.4419,  ..., -1.4273, -1.3689, -1.4711],
         [-1.4857, -1.5587, -1.5295,  ..., -1.4419, -1.4565, -1.5003],
         ...,
         [ 0.9814,  1.2150,  1.1420,  ...,  1.0398,  0.7333,  0.5727],
         [ 0.8938,  0.9376,  1.0982,  ...,  0.9960,  0.4121,  0.1055],
         [ 0.8647,  0.7479,  0.7479,  ...,  0.6165,  0.6895,  0.3537]],

        [[-1.4669, -1.3019, -1.1668,  ..., -1.4069, -1.3319, -1.4219],
         [-1.0467, -1.2268, -1.3919,  ..., -1.3769, -1.3169, -1.4219],
         [-1.4369, -1.5120, -1.4820,  ..., -1.3919, -1.4069, -1.4519],
         ...,
         [ 1.0994,  1.3395,  1.2645,  ...,  1.1594,  0.8442,  0.6792],
         [ 1.0093,  1.0544,  1.2194,  ...,  1.1144,  0.5141,  0.1989],
         [ 0.9793,  0.8593,  0.8593,  ...,  0.7242,  0.7992,  0.4540]],

        [[-1.2100, -1.0536, -0.9256,  ..., -1.1532, -1.0821, -1.1674],
         [-0.8119, -0.9825, -1.1389,  ..., -1.1247, -1.0678, -1.1674],
         [-1.1816, -1.2527, -1.2243,  ..., -1.1389, -1.1532, -1.1958],
         ...,
         [ 1.2216,  1.4491,  1.3780,  ...,  1.2785,  0.9799,  0.8234],
         [ 1.1363,  1.1789,  1.3354,  ...,  1.2358,  0.6670,  0.3684],
         [ 1.1078,  0.9941,  0.9941,  ...,  0.8661,  0.9372,  0.6101]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the a child skiing that is being helped by a man in this image? ASSISTANT: This view focuses on the a child skiing that is being helped by a man.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the young skier being held up by instructor shown in this image. ASSISTANT: Displayed here is the young skier being held up by instructor.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the black mens skis on the white snow in this picture? ASSISTANT: This is the selected black mens skis on the white snow.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (676, 1024), ['<image>\nCan you segment the a child skiing that is being helped by a man in this image?', '<image>\nDisplay a segmentation mask for the young skier being held up by instructor shown in this image.', '<image>\nCould you isolate the black mens skis on the white snow in this picture?'], ['a child skiing that is being helped by a man', 'young skier being held up by instructor', 'black mens skis on the white snow'], [{'a child skiing that is being helped by a man': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'young skier being held up by instructor': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'black mens skis on the white snow': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000472581.jpg', tensor([[[-1.2617, -1.2274, -1.1932,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2445, -1.2103, -1.1760,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2103, -1.1932, -1.1589,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-2.0665, -2.0665, -2.0665,  ...,  0.0000,  0.0000,  0.0000],
         [-2.0494, -2.0494, -2.0494,  ...,  0.0000,  0.0000,  0.0000],
         [-2.0494, -2.0494, -2.0494,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.3354, -1.3004, -1.2654,  ...,  0.0000,  0.0000,  0.0000],
         [-1.3354, -1.3004, -1.2479,  ...,  0.0000,  0.0000,  0.0000],
         [-1.3179, -1.2829, -1.2304,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.9832, -1.9832, -1.9832,  ...,  0.0000,  0.0000,  0.0000],
         [-1.9657, -1.9657, -1.9657,  ...,  0.0000,  0.0000,  0.0000],
         [-1.9657, -1.9657, -1.9657,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.1247, -1.1596, -1.1944,  ...,  0.0000,  0.0000,  0.0000],
         [-1.1073, -1.1421, -1.1770,  ...,  0.0000,  0.0000,  0.0000],
         [-1.0898, -1.1247, -1.1596,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.7522, -1.7522, -1.7522,  ...,  0.0000,  0.0000,  0.0000],
         [-1.7347, -1.7347, -1.7347,  ...,  0.0000,  0.0000,  0.0000],
         [-1.7347, -1.7347, -1.7347,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.7266, -0.7266, -0.6974,  ..., -0.4784, -0.5222, -0.4930],
         [-0.8142, -0.6974, -0.6682,  ..., -0.4930, -0.5076, -0.4492],
         [-0.7996, -0.6974, -0.6536,  ..., -0.4492, -0.4784, -0.4638],
         ...,
         [-0.7704, -0.3908, -0.5368,  ..., -1.3543, -1.3835, -1.3981],
         [-0.8726, -0.3908, -0.5076,  ..., -1.3251, -1.3981, -1.3835],
         [-1.0039, -0.5952, -0.4346,  ..., -1.3543, -1.3689, -1.3835]],

        [[-0.8816, -0.8216, -0.8216,  ..., -0.5065, -0.5065, -0.5665],
         [-0.8366, -0.8366, -0.8066,  ..., -0.5065, -0.5215, -0.5515],
         [-0.8216, -0.8516, -0.8066,  ..., -0.4914, -0.5065, -0.5515],
         ...,
         [-0.9417, -0.5815, -0.6565,  ..., -1.5120, -1.5420, -1.5420],
         [-1.0017, -0.6415, -0.6115,  ..., -1.4970, -1.5270, -1.5120],
         [-1.1368, -0.8516, -0.5515,  ..., -1.4970, -1.5120, -1.5120]],

        [[-0.7692, -0.7123, -0.6981,  ..., -0.3568, -0.4137, -0.4422],
         [-0.7834, -0.7266, -0.6697,  ..., -0.3568, -0.4137, -0.4279],
         [-0.7834, -0.7408, -0.6555,  ..., -0.3284, -0.3995, -0.4422],
         ...,
         [-0.7408, -0.4564, -0.5417,  ..., -1.3665, -1.3949, -1.3949],
         [-0.7977, -0.4279, -0.5417,  ..., -1.3522, -1.3949, -1.3807],
         [-0.9541, -0.6128, -0.4706,  ..., -1.3665, -1.3807, -1.3807]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the primary objects in the room? ASSISTANT: The primary objects in the room are a TV, a bookshelf filled with books, and a bed.</s>USER: What type of television is in the room, and where is it positioned? ASSISTANT: The room has a small black television with long antennas, positioned on a table near the bed and bookshelf.</s>USER: What is the color of the bookshelf in the room? ASSISTANT: The color of the bookshelf in the room is black.</s>USER: Can you describe the overall setting and ambiance of the room? ASSISTANT: Based on the image, the room appears to be a cozy and intimate space, likely serving as a bedroom or a personal retreat for relaxation and leisure. The presence of a bed implies that it is primarily used for rest and sleep, while the bookshelf filled with books suggests that reading is a favored pastime for the room's occupant. The small black television with long antennas adds an element of entertainment, possibly for watching shows, movies, or news before bedtime or during leisure time. \n\nThe layout of the room, with the television sitting next to the black bookcase and close to the bed, further emphasizes the practicality and functionality of the space, allowing easy accessibility to various items and activities without taking up too much area. The assortment of books on the bookshelf showcases a variety of interests and possibly a love for reading, lending an intellectual and personal touch to the room. Overall, the setting evokes a comfortable and inviting atmosphere, perfect for unwinding and spending some quality time alone or engaging in quiet activities like reading or watching TV.</s>"], tensor([], size=(0, 640, 480)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 768), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the primary objects in the room? ASSISTANT: The primary objects in the room are a TV, a bookshelf filled with books, and a bed.</s>USER: What type of television is in the room, and where is it positioned? ASSISTANT: The room has a small black television with long antennas, positioned on a table near the bed and bookshelf.</s>USER: What is the color of the bookshelf in the room? ASSISTANT: The color of the bookshelf in the room is black.</s>USER: Can you describe the overall setting and ambiance of the room? ASSISTANT: Based on the image, the room appears to be a cozy and intimate space, likely serving as a bedroom or a personal retreat for relaxation and leisure. The presence of a bed implies that it is primarily used for rest and sleep, while the bookshelf filled with books suggests that reading is a favored pastime for the room's occupant. The small black television with long antennas adds an element of entertainment, possibly for watching shows, movies, or news before bedtime or during leisure time. \n\nThe layout of the room, with the television sitting next to the black bookcase and close to the bed, further emphasizes the practicality and functionality of the space, allowing easy accessibility to various items and activities without taking up too much area. The assortment of books on the bookshelf showcases a variety of interests and possibly a love for reading, lending an intellectual and personal touch to the room. Overall, the setting evokes a comfortable and inviting atmosphere, perfect for unwinding and spending some quality time alone or engaging in quiet activities like reading or watching TV.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the primary objects in the room? ASSISTANT: The primary objects in the room are a TV, a bookshelf filled with books, and a bed.</s>USER: What type of television is in the room, and where is it positioned? ASSISTANT: The room has a small black television with long antennas, positioned on a table near the bed and bookshelf.</s>USER: What is the color of the bookshelf in the room? ASSISTANT: The color of the bookshelf in the room is black.</s>USER: Can you describe the overall setting and ambiance of the room? ASSISTANT: Based on the image, the room appears to be a cozy and intimate space, likely serving as a bedroom or a personal retreat for relaxation and leisure. The presence of a bed implies that it is primarily used for rest and sleep, while the bookshelf filled with books suggests that reading is a favored pastime for the room's occupant. The small black television with long antennas adds an element of entertainment, possibly for watching shows, movies, or news before bedtime or during leisure time. \n\nThe layout of the room, with the television sitting next to the black bookcase and close to the bed, further emphasizes the practicality and functionality of the space, allowing easy accessibility to various items and activities without taking up too much area. The assortment of books on the bookshelf showcases a variety of interests and possibly a love for reading, lending an intellectual and personal touch to the room. Overall, the setting evokes a comfortable and inviting atmosphere, perfect for unwinding and spending some quality time alone or engaging in quiet activities like reading or watching TV.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/saiapr_tc-12/11/images/11207.jpg', tensor([[[-1.4672, -1.4500, -1.4158,  ..., -0.4911, -0.5424, -0.5596],
         [-1.4500, -1.4329, -1.3987,  ..., -0.4739, -0.5253, -0.5424],
         [-1.4158, -1.3987, -1.3644,  ..., -0.4568, -0.4911, -0.5082],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.7731, -1.7556, -1.7206,  ..., -1.0378, -1.0903, -1.1078],
         [-1.7556, -1.7381, -1.7031,  ..., -1.0203, -1.0728, -1.0903],
         [-1.7206, -1.7031, -1.6681,  ..., -1.0028, -1.0378, -1.0553],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.6476, -1.6302, -1.5953,  ..., -1.1770, -1.2293, -1.2467],
         [-1.6302, -1.6127, -1.5779,  ..., -1.1770, -1.2293, -1.2467],
         [-1.5953, -1.5779, -1.5430,  ..., -1.1596, -1.2119, -1.2293],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.3908, -0.2302, -0.3324,  ...,  0.6311,  0.5435,  0.4559],
         [-0.6390, -0.4200, -0.5222,  ...,  0.5873,  0.5143,  0.4851],
         [-0.6828, -0.4784, -0.5514,  ...,  0.6019,  0.5289,  0.5289],
         ...,
         [ 0.2077,  0.2077,  0.2515,  ..., -0.2886, -0.2740, -0.2594],
         [ 0.1931,  0.2515,  0.2953,  ..., -0.2886, -0.2740, -0.2740],
         [ 0.2369,  0.2515,  0.2369,  ..., -0.2886, -0.2740, -0.2740]],

        [[-1.0767, -1.0617, -1.0767,  ...,  0.1389,  0.0488, -0.0112],
         [-1.3319, -1.2568, -1.2869,  ...,  0.1239,  0.0638,  0.0338],
         [-1.3769, -1.3019, -1.2869,  ...,  0.1689,  0.1239,  0.1089],
         ...,
         [ 0.0338,  0.0338,  0.0638,  ..., -0.4014, -0.3864, -0.3714],
         [ 0.0188,  0.0789,  0.1239,  ..., -0.3864, -0.3864, -0.3864],
         [ 0.0638,  0.0789,  0.0638,  ..., -0.3864, -0.3864, -0.3864]],

        [[-1.1389, -1.1532, -1.1389,  ..., -0.4137, -0.4706, -0.5275],
         [-1.3949, -1.3522, -1.3665,  ..., -0.3995, -0.4564, -0.4706],
         [-1.4376, -1.4091, -1.4233,  ..., -0.3284, -0.3711, -0.3568],
         ...,
         [-0.0724, -0.0724, -0.0440,  ..., -0.4279, -0.4137, -0.3995],
         [-0.0867, -0.0298,  0.0129,  ..., -0.4279, -0.4137, -0.4137],
         [-0.0440, -0.0298, -0.0440,  ..., -0.4279, -0.4137, -0.4137]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the table from the image below? ASSISTANT: The table portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the white tablecloth on right of pic shown in this image. ASSISTANT: Displayed here is the white tablecloth on right of pic.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the any person in this image? ASSISTANT: Below you can see the any person.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the guy right from the provided image. ASSISTANT: Take a look at the guy right here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is guy far left in this image? Please elaborate your answer and explain why. ASSISTANT: Here's the area for guy far left.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nWould you please extract the table from the image below?', '<image>\nDisplay a segmentation mask for the white tablecloth on right of pic shown in this image.', '<image>\nCan you segment the any person in this image?', '<image>\nSegment the guy right from the provided image.', '<image>\nWhat is guy far left in this image? Please elaborate your answer and explain why.'], ['table', 'white tablecloth on right of pic', 'any person', 'guy right', 'guy far left'], [{'table': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'white tablecloth on right of pic': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'any person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'guy right': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'guy far left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/saiapr_tc-12/23/images/23747.jpg', tensor([[[2.0263, 2.0263, 2.0263,  ..., 2.1804, 2.1804, 2.1804],
         [2.0263, 2.0263, 2.0263,  ..., 2.1804, 2.1804, 2.1804],
         [2.0263, 2.0263, 2.0263,  ..., 2.1804, 2.1804, 2.1804],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.2185, 2.2185, 2.2185,  ..., 2.3761, 2.3761, 2.3761],
         [2.2185, 2.2185, 2.2185,  ..., 2.3761, 2.3761, 2.3761],
         [2.2185, 2.2185, 2.2185,  ..., 2.3761, 2.3761, 2.3761],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.5180, 2.5180, 2.5180,  ..., 2.6226, 2.6226, 2.6226],
         [2.5180, 2.5180, 2.5180,  ..., 2.6226, 2.6226, 2.6226],
         [2.5180, 2.5180, 2.5180,  ..., 2.6226, 2.6226, 2.6226],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.7698e+00,  1.7698e+00,  1.7698e+00,  ...,  1.9157e+00,
           1.9157e+00,  1.9157e+00],
         [ 1.7698e+00,  1.7698e+00,  1.7698e+00,  ...,  1.9157e+00,
           1.9157e+00,  1.9157e+00],
         [ 1.7844e+00,  1.7844e+00,  1.7844e+00,  ...,  1.9157e+00,
           1.9157e+00,  1.9157e+00],
         ...,
         [-6.9648e-02,  3.2541e-02,  1.7853e-01,  ..., -2.5853e-02,
          -1.1255e-02, -1.1255e-02],
         [ 6.1738e-02,  1.3473e-01,  1.0553e-01,  ...,  3.3439e-03,
           1.7942e-02,  7.6336e-02],
         [-4.0451e-02, -6.9648e-02, -1.2804e-01,  ...,  9.0935e-02,
           1.2013e-01,  1.6393e-01]],

        [[ 1.9098e+00,  1.9098e+00,  1.9248e+00,  ...,  2.0299e+00,
           2.0299e+00,  2.0299e+00],
         [ 1.9098e+00,  1.9098e+00,  1.9248e+00,  ...,  2.0299e+00,
           2.0299e+00,  2.0299e+00],
         [ 1.9248e+00,  1.9248e+00,  1.9248e+00,  ...,  2.0299e+00,
           2.0299e+00,  2.0299e+00],
         ...,
         [-2.3631e-01, -1.3126e-01,  1.8820e-02,  ..., -2.8134e-01,
          -2.5132e-01, -2.6633e-01],
         [-1.1625e-01, -4.1212e-02, -5.6219e-02,  ..., -2.2130e-01,
          -2.0630e-01, -1.7628e-01],
         [-2.3631e-01, -2.6633e-01, -3.2636e-01,  ..., -1.4627e-01,
          -8.6235e-02, -5.6219e-02]],

        [[ 2.1032e+00,  2.1032e+00,  2.1032e+00,  ...,  2.1459e+00,
           2.1459e+00,  2.1459e+00],
         [ 2.1032e+00,  2.1032e+00,  2.1032e+00,  ...,  2.1459e+00,
           2.1459e+00,  2.1459e+00],
         [ 2.1175e+00,  2.1175e+00,  2.1032e+00,  ...,  2.1459e+00,
           2.1459e+00,  2.1459e+00],
         ...,
         [-2.1463e-01, -1.1509e-01, -1.3329e-03,  ..., -3.5683e-01,
          -3.4261e-01, -3.4261e-01],
         [-1.0087e-01, -4.3993e-02, -8.6653e-02,  ..., -2.8573e-01,
          -2.8573e-01, -2.4307e-01],
         [-2.4307e-01, -2.7151e-01, -3.2839e-01,  ..., -2.1463e-01,
          -1.8619e-01, -1.4353e-01]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the van bus thing anywhere from the provided image. ASSISTANT: Displayed here is the van bus thing anywhere.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the sky in this picture. ASSISTANT: This view focuses on the sky.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the the floor under truck in this picture? ASSISTANT: Here's where the the floor under truck appears in the image.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nSegment the van bus thing anywhere from the provided image.', '<image>\nPlease point out the sky in this picture.', '<image>\nCould you isolate the the floor under truck in this picture?'], ['van bus thing anywhere', 'sky', 'the floor under truck'], [{'van bus thing anywhere': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'sky': tensor([[1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the floor under truck': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/saiapr_tc-12/17/images/17506.jpg', tensor([[[-1.6898, -1.6898, -1.7069,  ..., -2.0837, -2.0837, -2.0837],
         [-1.6898, -1.6898, -1.7069,  ..., -2.0837, -2.0837, -2.0837],
         [-1.6898, -1.6898, -1.7069,  ..., -2.0837, -2.0837, -2.0837],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.7206, -1.7206, -1.7381,  ..., -2.0007, -2.0007, -2.0007],
         [-1.7206, -1.7206, -1.7381,  ..., -2.0007, -2.0007, -2.0007],
         [-1.7206, -1.7206, -1.7381,  ..., -2.0007, -2.0007, -2.0007],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.5953, -1.5953, -1.6127,  ..., -1.7696, -1.7696, -1.7696],
         [-1.5953, -1.5953, -1.6127,  ..., -1.7696, -1.7696, -1.7696],
         [-1.5953, -1.5953, -1.6127,  ..., -1.7696, -1.7696, -1.7696],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.4127, -1.4273, -1.4565,  ..., -1.5733, -1.5295, -1.5149],
         [-1.4127, -1.4273, -1.4565,  ..., -1.6025, -1.5441, -1.5149],
         [-1.4127, -1.4273, -1.4565,  ..., -1.6317, -1.5733, -1.5295],
         ...,
         [ 0.3975,  0.3537,  0.2661,  ..., -0.7850, -0.9164, -1.0039],
         [ 0.3099,  0.0179,  0.1201,  ..., -0.9310, -1.0331, -1.0331],
         [ 0.7187, -0.0259, -0.0405,  ..., -0.7120, -0.9456, -0.9602]],

        [[-1.4669, -1.4820, -1.5120,  ..., -1.5720, -1.5420, -1.5270],
         [-1.4669, -1.4820, -1.5120,  ..., -1.6020, -1.5570, -1.5270],
         [-1.4669, -1.4820, -1.5120,  ..., -1.6320, -1.5870, -1.5420],
         ...,
         [-0.2813, -0.2213, -0.3864,  ..., -1.1668, -1.3019, -1.4219],
         [-0.4464, -0.6265, -0.5815,  ..., -1.2869, -1.3769, -1.4069],
         [-0.1012, -0.7316, -0.7616,  ..., -1.1068, -1.2268, -1.3019]],

        [[-1.2954, -1.3096, -1.3380,  ..., -1.3380, -1.3238, -1.3096],
         [-1.2954, -1.3096, -1.3380,  ..., -1.3522, -1.3380, -1.3096],
         [-1.2954, -1.3096, -1.3380,  ..., -1.3807, -1.3665, -1.3238],
         ...,
         [-0.0582, -0.0440, -0.1293,  ..., -1.1247, -1.2527, -1.3807],
         [-0.2004, -0.4137, -0.3000,  ..., -1.1674, -1.2669, -1.3380],
         [ 0.1409, -0.4848, -0.4564,  ..., -0.9256, -1.0963, -1.2100]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the person in this image. ASSISTANT: I've marked the person for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the flowers from the image below? ASSISTANT: The region corresponding to the flowers is shown.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nPlease segment the person in this image.', '<image>\nWould you please extract the flowers from the image below?'], ['person', 'flowers'], [{'person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'flowers': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000424582.jpg', tensor([[[-0.6452, -0.6452, -0.6281,  ..., -0.6109, -0.6452, -0.6623],
         [-0.6452, -0.6452, -0.6452,  ..., -0.5082, -0.5596, -0.5767],
         [-0.6623, -0.6623, -0.6623,  ..., -0.3198, -0.3883, -0.4226],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.3704, -1.3704, -1.3529,  ..., -1.4055, -1.4580, -1.4755],
         [-1.3704, -1.3880, -1.3704,  ..., -1.2829, -1.3529, -1.3704],
         [-1.3880, -1.4055, -1.4230,  ..., -1.0553, -1.1254, -1.1604],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.5256, -1.5256, -1.5081,  ..., -1.5430, -1.5604, -1.5604],
         [-1.5604, -1.5604, -1.5604,  ..., -1.4036, -1.4384, -1.4384],
         [-1.6476, -1.6476, -1.6476,  ..., -1.1073, -1.1770, -1.1944],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.3178, -0.3178, -0.3178,  ...,  0.0617,  0.0325,  0.1201],
         [-0.4346, -0.4492, -0.3908,  ..., -0.1134, -0.0696, -0.1572],
         [-0.0405, -0.1572, -0.2010,  ..., -0.2302, -0.1426, -0.2010],
         ...,
         [-1.5587, -1.5441, -1.5441,  ..., -0.5368, -0.5806, -0.5660],
         [-1.6171, -1.5879, -1.5733,  ..., -0.5806, -0.6098, -0.5660],
         [-1.6025, -1.6171, -1.6025,  ..., -0.5514, -0.6682, -0.6536]],

        [[-1.1068, -1.1068, -1.1218,  ..., -0.8366, -0.8666, -0.7316],
         [-1.2418, -1.3019, -1.2118,  ..., -1.0617, -0.9867, -1.0467],
         [-0.8516, -0.9267, -0.9867,  ..., -1.1068, -1.0617, -1.0918],
         ...,
         [-1.5270, -1.5120, -1.5120,  ..., -0.3114, -0.3264, -0.3114],
         [-1.5870, -1.5570, -1.5270,  ..., -0.3414, -0.3564, -0.3114],
         [-1.5870, -1.6020, -1.5420,  ..., -0.3264, -0.4164, -0.4164]],

        [[-1.1389, -1.1532, -1.1674,  ..., -1.1389, -1.1532, -1.0252],
         [-1.3522, -1.3665, -1.3807,  ..., -1.2527, -1.1674, -1.2385],
         [-0.9825, -1.0678, -1.0821,  ..., -1.3096, -1.3380, -1.3665],
         ...,
         [-1.1674, -1.1389, -1.1816,  ...,  0.0555,  0.0413,  0.0698],
         [-1.1958, -1.1816, -1.1958,  ...,  0.0271,  0.0129,  0.0840],
         [-1.1674, -1.2243, -1.2243,  ...,  0.0413, -0.0440, -0.0440]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the dining table in this image? ASSISTANT: Below you can see the dining table.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the broccoli in this picture? ASSISTANT: I've marked the broccoli for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the carrot shown in this image. ASSISTANT: You can see the carrot in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the table in this picture? ASSISTANT: The table is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the fork in this image? ASSISTANT: Below you can see the fork.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ...,  True,  True,  True],
         [False, False, False,  ...,  True,  True,  True],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         ...,
         [False, False, False,  ...,  True,  True,  True],
         [False, False, False,  ...,  True,  True,  True],
         [False, False, False,  ...,  True,  True,  True]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[164, 164, 164,  ..., 164, 164, 164],
        [164, 164, 164,  ...,  66,  66,  66],
        [164, 164, 164,  ...,  66,  66,  66],
        ...,
        [255, 255, 255,  ..., 164, 164, 164],
        [255, 255, 255,  ..., 164, 164, 164],
        [255, 255, 255,  ..., 164, 164, 164]]), (768, 1024), ['<image>\nCan you segment the dining table in this image?', '<image>\nCould you isolate the broccoli in this picture?', '<image>\nDisplay a segmentation mask for the carrot shown in this image.', '<image>\nCould you isolate the table in this picture?', '<image>\nCan you segment the fork in this image?'], ['dining table', 'broccoli', 'carrot', 'table', 'fork'], [{'dining table': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ...,  True,  True,  True],
        [False, False, False,  ...,  True,  True,  True],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'broccoli': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'carrot': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'table': tensor([[ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        ...,
        [False, False, False,  ...,  True,  True,  True],
        [False, False, False,  ...,  True,  True,  True],
        [False, False, False,  ...,  True,  True,  True]])}, {'fork': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg')]
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000133384.jpg', tensor([[[ 0.7762,  0.7762,  0.7762,  ..., -1.7925, -1.7925, -1.7925],
         [ 0.7591,  0.7762,  0.7762,  ..., -1.7412, -1.7583, -1.7583],
         [ 0.7419,  0.7591,  0.7762,  ..., -1.6898, -1.7069, -1.7240],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.5903,  0.5903,  0.5903,  ..., -1.7206, -1.7206, -1.7206],
         [ 0.5903,  0.5903,  0.5903,  ..., -1.7031, -1.6856, -1.6856],
         [ 0.5728,  0.5728,  0.5728,  ..., -1.6681, -1.6506, -1.6506],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.2871,  0.2871,  0.2871,  ..., -1.5604, -1.5604, -1.5604],
         [ 0.3219,  0.3219,  0.3219,  ..., -1.5256, -1.5256, -1.5256],
         [ 0.3742,  0.3742,  0.3742,  ..., -1.4907, -1.4907, -1.4907],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.0252,  1.0252,  1.0398,  ..., -0.7266, -0.6828, -0.7412],
         [ 1.0544,  1.0398,  1.0544,  ..., -0.7120, -0.6828, -0.7266],
         [ 1.0398,  1.0690,  1.0690,  ..., -0.7120, -0.6682, -0.7558],
         ...,
         [ 0.9522,  0.9522,  0.9814,  ...,  0.9960,  0.9814,  1.0836],
         [ 0.9376,  0.9668,  0.9668,  ...,  0.7625,  0.8209,  0.8209],
         [ 0.9084,  0.8938,  0.9376,  ...,  0.6749,  0.7479,  0.7625]],

        [[ 0.8743,  0.8893,  0.9343,  ..., -0.7916, -0.8216, -0.8066],
         [ 0.9193,  0.9343,  0.9493,  ..., -0.7766, -0.8066, -0.7916],
         [ 0.9343,  0.9643,  0.9493,  ..., -0.7766, -0.7766, -0.7916],
         ...,
         [-0.3114, -0.3264, -0.2963,  ...,  0.4090,  0.4090,  0.5891],
         [-0.3264, -0.3264, -0.3114,  ...,  0.2589,  0.3490,  0.3490],
         [-0.3264, -0.3714, -0.3114,  ...,  0.1539,  0.2439,  0.2890]],

        [[ 0.7381,  0.7239,  0.7097,  ..., -0.8403, -0.7834, -0.7408],
         [ 0.6955,  0.6812,  0.6955,  ..., -0.7550, -0.7692, -0.7834],
         [ 0.6955,  0.7239,  0.7239,  ..., -0.7692, -0.7692, -0.8119],
         ...,
         [ 0.1266,  0.1551,  0.1551,  ...,  0.3826,  0.3826,  0.5390],
         [ 0.1124,  0.1835,  0.1409,  ...,  0.2404,  0.3115,  0.3257],
         [ 0.1835,  0.1977,  0.1693,  ...,  0.1266,  0.2120,  0.2404]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the left man in this picture? ASSISTANT: This view focuses on the left man.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the man on the right in this picture? ASSISTANT: This is the selected man on the right.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the front guy in this picture? ASSISTANT: The front guy portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the dude on the right in this image? ASSISTANT: Sure — this shows the dude on the right.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the left guy in this picture? ASSISTANT: The left guy portion is displayed below.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nCould you isolate the left man in this picture?', '<image>\nCould you identify and segment out the man on the right in this picture?', '<image>\nCould you identify and segment out the front guy in this picture?', '<image>\nCan you segment the dude on the right in this image?', '<image>\nCould you isolate the left guy in this picture?'], ['left man', 'man on the right', 'front guy', 'dude on the right', 'left guy'], [{'left man': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'man on the right': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'front guy': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'dude on the right': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'left guy': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000068601.jpg', tensor([[[0.1768, 0.2111, 0.2624,  ..., 0.5878, 0.5707, 0.5707],
         [0.1597, 0.1939, 0.2453,  ..., 0.5878, 0.5707, 0.5707],
         [0.1426, 0.1768, 0.2282,  ..., 0.5707, 0.5707, 0.5707],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.5028, 0.5203, 0.5553,  ..., 0.7654, 0.7479, 0.7479],
         [0.4853, 0.5028, 0.5378,  ..., 0.7654, 0.7479, 0.7479],
         [0.4503, 0.4678, 0.5028,  ..., 0.7479, 0.7479, 0.7479],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.6531, 0.7054, 0.7576,  ..., 0.8971, 0.8797, 0.8797],
         [0.6531, 0.7054, 0.7576,  ..., 0.8971, 0.8797, 0.8797],
         [0.6531, 0.7054, 0.7576,  ..., 0.8797, 0.8797, 0.8797],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 0.7917,  0.8063,  0.7625,  ...,  0.8063,  0.8647,  0.8792],
         [ 0.7917,  0.8063,  0.7771,  ...,  0.8209,  0.8792,  0.9084],
         [ 0.8063,  0.8063,  0.7917,  ...,  0.8355,  0.8792,  0.9084],
         ...,
         [ 0.2077,  0.1347,  0.4267,  ..., -0.1572, -0.1426, -0.2010],
         [ 0.2807,  0.0471,  0.3245,  ..., -0.1572, -0.1134, -0.1572],
         [ 0.3975,  0.1055,  0.2515,  ..., -0.1864, -0.1280, -0.1572]],

        [[ 0.9343,  0.9343,  0.9193,  ...,  0.9493,  1.0093,  1.0243],
         [ 0.9343,  0.9493,  0.9343,  ...,  0.9643,  1.0243,  1.0544],
         [ 0.9343,  0.9493,  0.9493,  ...,  0.9793,  1.0243,  1.0544],
         ...,
         [ 0.3790,  0.2589,  0.5891,  ..., -0.1163, -0.0712, -0.0562],
         [ 0.4991,  0.2589,  0.5291,  ..., -0.1313, -0.0862, -0.0562],
         [ 0.5741,  0.2890,  0.3790,  ..., -0.1012, -0.0712, -0.0712]],

        [[ 0.8945,  0.9514,  0.9514,  ...,  1.0225,  1.0936,  1.0936],
         [ 0.9230,  0.9656,  0.9656,  ...,  1.0367,  1.1078,  1.1363],
         [ 0.9514,  0.9514,  0.9372,  ...,  1.0367,  1.1078,  1.1363],
         ...,
         [ 0.5959,  0.6670,  0.9514,  ...,  0.0129,  0.0840,  0.0555],
         [ 0.7239,  0.5532,  0.9230,  ...,  0.0129,  0.0982,  0.0840],
         [ 0.8377,  0.5390,  0.8234,  ...,  0.0555,  0.1551,  0.1124]]]), ['A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human\'s questions. USER: <image>\nWhat are the concerns about the design and condition of this bathroom? ASSISTANT: The bathroom presents several concerns in terms of both its design and condition. The two toilets are positioned adjacently without any privacy partitions, which makes for an awkward, uncomfortable, and potentially unsanitary experience for users. The exposed piping beneath the toilets may not only be visually unappealing but could also indicate that the plumbing is old or not properly maintained.\n\nAdditionally, the general description of the bathroom as "run-down" and "dingy" suggests that the area has not been well-maintained, raising concerns about cleanliness and hygiene. One of the toilets is missing its lid, further emphasizing the neglect in maintenance. Also, there is a toilet brush placed in one of the toilets, which is quite unusual and unhygienic. These concerns might discourage people from using the bathroom or potentially pose health risks. Issues like these emphasize the importance of proper bathroom design and maintenance to ensure a functional, sanitary, and comfortable space.</s>'], tensor([], size=(0, 480, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human\'s questions. USER: <image>\nWhat are the concerns about the design and condition of this bathroom? ASSISTANT: The bathroom presents several concerns in terms of both its design and condition. The two toilets are positioned adjacently without any privacy partitions, which makes for an awkward, uncomfortable, and potentially unsanitary experience for users. The exposed piping beneath the toilets may not only be visually unappealing but could also indicate that the plumbing is old or not properly maintained.\n\nAdditionally, the general description of the bathroom as "run-down" and "dingy" suggests that the area has not been well-maintained, raising concerns about cleanliness and hygiene. One of the toilets is missing its lid, further emphasizing the neglect in maintenance. Also, there is a toilet brush placed in one of the toilets, which is quite unusual and unhygienic. These concerns might discourage people from using the bathroom or potentially pose health risks. Issues like these emphasize the importance of proper bathroom design and maintenance to ensure a functional, sanitary, and comfortable space.</s>'], ['A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human\'s questions. USER: <image>\nWhat are the concerns about the design and condition of this bathroom? ASSISTANT: The bathroom presents several concerns in terms of both its design and condition. The two toilets are positioned adjacently without any privacy partitions, which makes for an awkward, uncomfortable, and potentially unsanitary experience for users. The exposed piping beneath the toilets may not only be visually unappealing but could also indicate that the plumbing is old or not properly maintained.\n\nAdditionally, the general description of the bathroom as "run-down" and "dingy" suggests that the area has not been well-maintained, raising concerns about cleanliness and hygiene. One of the toilets is missing its lid, further emphasizing the neglect in maintenance. Also, there is a toilet brush placed in one of the toilets, which is quite unusual and unhygienic. These concerns might discourage people from using the bathroom or potentially pose health risks. Issues like these emphasize the importance of proper bathroom design and maintenance to ensure a functional, sanitary, and comfortable space.</s>'], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/GKMwQR0Jw6rbZLUNtkfrVQ.jpg', tensor([[[-1.7240, -1.6384, -1.7412,  ...,  1.2214,  1.2214,  1.1872],
         [-1.6727, -1.6727, -1.6555,  ...,  1.2214,  1.2214,  1.1872],
         [-1.5870, -1.7240, -1.6898,  ...,  1.2214,  1.2043,  1.1872],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.5805, -1.4930, -1.5980,  ...,  1.4832,  1.4832,  1.4482],
         [-1.5280, -1.5280, -1.5105,  ...,  1.4832,  1.4832,  1.4482],
         [-1.4405, -1.5805, -1.5455,  ...,  1.4832,  1.4657,  1.4482],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.1944, -1.1073, -1.2119,  ...,  1.9080,  1.9080,  1.8731],
         [-1.1421, -1.1421, -1.1247,  ...,  1.9080,  1.9080,  1.8731],
         [-1.0550, -1.1944, -1.1596,  ...,  1.9080,  1.8905,  1.8731],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.1566,  1.1566,  1.1566,  ...,  1.1858,  1.2004,  1.2004],
         [ 1.1566,  1.1566,  1.1420,  ...,  1.2004,  1.2004,  1.1858],
         [ 1.1712,  1.1566,  1.1566,  ...,  1.2004,  1.2004,  1.2004],
         ...,
         [-1.1937, -1.0477, -1.0769,  ..., -0.7266, -0.7266, -0.7120],
         [-1.0331, -1.1499, -1.1353,  ..., -0.7996, -0.7850, -0.7412],
         [-1.0623, -1.2521, -1.3251,  ..., -0.8580, -0.8434, -0.8288]],

        [[ 1.3395,  1.3395,  1.3545,  ...,  1.3695,  1.3845,  1.3695],
         [ 1.3395,  1.3395,  1.3395,  ...,  1.3845,  1.3845,  1.3695],
         [ 1.3545,  1.3395,  1.3395,  ...,  1.3845,  1.3845,  1.3845],
         ...,
         [-1.1518, -1.0317, -1.0767,  ..., -0.6565, -0.6565, -0.6415],
         [-1.0467, -1.1368, -1.1218,  ..., -0.7166, -0.7166, -0.6715],
         [-1.1218, -1.2418, -1.2718,  ..., -0.7916, -0.7766, -0.7466]],

        [[ 1.6055,  1.6198,  1.6340,  ...,  1.6482,  1.6624,  1.6482],
         [ 1.6198,  1.6198,  1.6198,  ...,  1.6624,  1.6624,  1.6482],
         [ 1.6340,  1.6198,  1.6198,  ...,  1.6624,  1.6624,  1.6624],
         ...,
         [-0.9399, -0.8688, -0.9114,  ..., -0.4137, -0.4137, -0.4279],
         [-0.9256, -1.0252, -1.0110,  ..., -0.4848, -0.4848, -0.4564],
         [-1.0252, -1.1247, -1.0678,  ..., -0.5844, -0.5559, -0.5275]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the traffic sign - parking in this image. ASSISTANT: Below you can see the traffic sign - parking.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the traffic island in this picture. ASSISTANT: I've marked the traffic island for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the parking shown in this image. ASSISTANT: Here's where the parking appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the curb cut in this picture? ASSISTANT: Displayed here is the curb cut.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the traffic light - general (upright) from the provided image. ASSISTANT: Displayed here is the traffic light - general (upright).</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        ...,
        [25, 25, 25,  ..., 21, 21, 21],
        [25, 25, 25,  ..., 21, 21, 21],
        [25, 25, 25,  ..., 21, 21, 21]]), (768, 1024), ['<image>\nPlease highlight the traffic sign - parking in this image.', '<image>\nPlease point out the traffic island in this picture.', '<image>\nDisplay a segmentation mask for the parking shown in this image.', '<image>\nCould you identify and segment out the curb cut in this picture?', '<image>\nSegment the traffic light - general (upright) from the provided image.'], ['traffic sign - parking', 'traffic island', 'parking', 'curb cut', 'traffic light - general (upright)'], [{'traffic sign - parking': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'traffic island': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False]])}, {'parking': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'curb cut': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'traffic light - general (upright)': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2010_003435.jpg', tensor([[[-2.0494, -2.0494, -2.0665,  ..., -1.9980, -1.9638, -1.9467],
         [-2.0494, -2.0494, -2.0665,  ..., -1.9980, -1.9809, -1.9638],
         [-2.0494, -2.0494, -2.0665,  ..., -2.0152, -1.9980, -1.9809],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.8431, -1.8431, -1.8606,  ..., -1.4580, -1.4055, -1.3704],
         [-1.8431, -1.8431, -1.8606,  ..., -1.5280, -1.4755, -1.4405],
         [-1.8431, -1.8431, -1.8606,  ..., -1.6681, -1.6155, -1.5805],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.7522, -1.7522, -1.7696,  ..., -1.5779, -1.5430, -1.5256],
         [-1.7522, -1.7522, -1.7696,  ..., -1.6127, -1.5779, -1.5604],
         [-1.7522, -1.7522, -1.7696,  ..., -1.6824, -1.6476, -1.6476],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.7923, -1.7485, -1.7339,  ..., -1.7631, -1.7631, -1.7631],
         [-1.7485, -1.7339, -1.7339,  ..., -1.7923, -1.7485, -1.6463],
         [-1.7193, -1.7339, -1.7339,  ..., -1.7631, -1.6609, -1.5733],
         ...,
         [ 1.9011,  1.9011,  1.9303,  ...,  1.2296,  1.2150,  1.2150],
         [ 1.9157,  1.9011,  1.9303,  ...,  1.1858,  1.1712,  1.1712],
         [ 1.9157,  1.9157,  1.9303,  ...,  1.1420,  1.1420,  1.1274]],

        [[-1.4970, -1.5120, -1.5270,  ..., -1.6320, -1.6771, -1.6170],
         [-1.5120, -1.5270, -1.5270,  ..., -1.7371, -1.6170, -1.3619],
         [-1.5270, -1.5270, -1.5270,  ..., -1.6320, -1.3919, -1.1518],
         ...,
         [ 1.9848,  2.0599,  2.0749,  ...,  1.3395,  1.3395,  1.3695],
         [ 1.9548,  2.0599,  2.0749,  ...,  1.2945,  1.3095,  1.3245],
         [ 1.9398,  2.0599,  2.0749,  ...,  1.2495,  1.2795,  1.2945]],

        [[-1.4518, -1.4518, -1.4660,  ..., -1.4802, -1.4660, -1.4660],
         [-1.4518, -1.4518, -1.4660,  ..., -1.4802, -1.4518, -1.4091],
         [-1.4518, -1.4660, -1.4660,  ..., -1.4518, -1.3949, -1.4233],
         ...,
         [ 1.9326,  2.0606,  2.0890,  ...,  1.5344,  1.5202,  1.5487],
         [ 1.8899,  2.0464,  2.0890,  ...,  1.4918,  1.4918,  1.5060],
         [ 1.8473,  2.0321,  2.0890,  ...,  1.4491,  1.4633,  1.4776]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is cat leg in this image? Please elaborate your answer and explain why. ASSISTANT: This is the selected cat leg.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the the eye of the cat in this image. ASSISTANT: The the eye of the cat is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the cat eye in this image. ASSISTANT: Take a look at the cat eye here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the the torso of the cat from the image below? ASSISTANT: I've marked the the torso of the cat for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the cat nose shown in this image. ASSISTANT: This is the selected cat nose.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [1, 1, 1,  ..., 0, 0, 0],
         [1, 1, 1,  ..., 0, 0, 0],
         [1, 1, 1,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 1, 1, 1],
         [0, 0, 0,  ..., 1, 1, 1],
         [0, 0, 0,  ..., 1, 1, 1]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nWhat is cat leg in this image? Please elaborate your answer and explain why.', '<image>\nPlease segment the the eye of the cat in this image.', '<image>\nPlease segment the cat eye in this image.', '<image>\nWould you please extract the the torso of the cat from the image below?', '<image>\nDisplay a segmentation mask for the cat nose shown in this image.'], ['cat leg', 'the eye of the cat', 'cat eye', 'the torso of the cat', 'cat nose'], [{'cat leg': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [1, 1, 1,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the eye of the cat': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'cat eye': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the torso of the cat': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 1, 1, 1],
        [0, 0, 0,  ..., 1, 1, 1],
        [0, 0, 0,  ..., 1, 1, 1]], dtype=torch.uint8)}, {'cat nose': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000406954.jpg', tensor([[[-2.1179, -1.8097, -1.3987,  ...,  1.7694,  2.1290,  2.2489],
         [-2.1008, -1.8610, -1.5357,  ...,  0.7762,  0.8961,  0.8961],
         [-2.0494, -1.9124, -1.7240,  ..., -0.5082, -0.6965, -0.8678],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.8782, -1.3880, -0.7402,  ...,  2.2360,  2.3936,  2.4111],
         [-1.8782, -1.5280, -1.0728,  ...,  1.3782,  1.3782,  1.3081],
         [-1.8606, -1.7206, -1.5280,  ...,  0.2402,  0.0301, -0.1450],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.0201, -0.8284, -0.5495,  ...,  0.9842,  1.5768,  1.9777],
         [-1.0550, -0.8807, -0.6541,  ...,  0.4962,  0.7925,  0.9842],
         [-1.0724, -0.9504, -0.7936,  ..., -0.1312, -0.2184, -0.2881],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.6171, -0.1718,  1.6092,  ...,  1.7990,  0.3829,  1.2880],
         [-1.6025,  0.5289,  1.7260,  ...,  1.7844,  1.4778,  0.5289],
         [-1.5441,  0.6165,  1.9303,  ...,  1.6530,  1.9303, -0.0259],
         ...,
         [-0.0988, -0.6682, -1.2667,  ...,  0.9814,  0.8938,  1.1566],
         [-0.4054,  0.0325, -0.2010,  ...,  0.2953,  0.6311,  0.5873],
         [-0.0696,  0.1055,  0.5581,  ..., -0.7266,  0.2807, -0.0405]],

        [[-1.3319,  0.0038,  1.6997,  ...,  1.8948,  0.5741,  1.5346],
         [-1.3319,  0.7392,  1.7747,  ...,  1.9248,  1.5946,  0.8893],
         [-1.2568,  0.8142,  2.0599,  ...,  1.9398,  2.0599,  0.4991],
         ...,
         [ 0.0939, -0.1913, -0.6265,  ...,  1.0393,  0.7542,  0.9943],
         [-0.3564, -0.1012, -0.3264,  ...,  0.4090,  0.6191,  0.6942],
         [-0.2813, -0.2213,  0.2439,  ..., -0.6115,  0.3490,  0.1389]],

        [[-0.7692,  0.0555,  1.1647,  ...,  1.6482,  0.3115,  1.0225],
         [-0.7408,  0.6101,  0.9799,  ...,  1.7193,  1.5060,  0.5817],
         [-0.7692,  0.6670,  1.4491,  ...,  1.9468,  1.8757,  0.2404],
         ...,
         [-0.8403, -0.8830, -0.9825,  ..., -0.4990, -0.4564, -0.3995],
         [-0.4564, -0.4279, -0.5559,  ..., -0.5559, -0.5844, -0.5844],
         [-0.4564, -0.4422, -0.3426,  ..., -0.7550, -0.5986, -0.7692]]]), ['A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human\'s questions. USER: <image>\nWhat are the consequences of a zebra being alone in its natural habitat? ASSISTANT: In its natural habitat, a zebra being alone can face several challenges and consequences. Zebras are social animals and typically live in groups called harems, which provide them with protection and improved chances of survival against predators such as lions, hyenas, and leopards. \n\nAn isolated zebra becomes more vulnerable to these predators as it lacks the advantage of "safety in numbers" that comes with being in a herd. In a group, zebras can collectively detect threats more effectively, and their unique striped patterns create an optical illusion known as "motion dazzle," which makes it harder for predators to target a specific individual during a chase.\n\nFurthermore, being alone can lead to increased stress and a lack of opportunities for social interaction, grooming, and reproduction, which are vital for the overall well-being and survival of zebras. Consequently, a lone zebra in the wild may have a harder time thriving and ensuring the continuation of its lineage.</s>'], tensor([], size=(0, 429, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (686, 1024), ['A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human\'s questions. USER: <image>\nWhat are the consequences of a zebra being alone in its natural habitat? ASSISTANT: In its natural habitat, a zebra being alone can face several challenges and consequences. Zebras are social animals and typically live in groups called harems, which provide them with protection and improved chances of survival against predators such as lions, hyenas, and leopards. \n\nAn isolated zebra becomes more vulnerable to these predators as it lacks the advantage of "safety in numbers" that comes with being in a herd. In a group, zebras can collectively detect threats more effectively, and their unique striped patterns create an optical illusion known as "motion dazzle," which makes it harder for predators to target a specific individual during a chase.\n\nFurthermore, being alone can lead to increased stress and a lack of opportunities for social interaction, grooming, and reproduction, which are vital for the overall well-being and survival of zebras. Consequently, a lone zebra in the wild may have a harder time thriving and ensuring the continuation of its lineage.</s>'], ['A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human\'s questions. USER: <image>\nWhat are the consequences of a zebra being alone in its natural habitat? ASSISTANT: In its natural habitat, a zebra being alone can face several challenges and consequences. Zebras are social animals and typically live in groups called harems, which provide them with protection and improved chances of survival against predators such as lions, hyenas, and leopards. \n\nAn isolated zebra becomes more vulnerable to these predators as it lacks the advantage of "safety in numbers" that comes with being in a herd. In a group, zebras can collectively detect threats more effectively, and their unique striped patterns create an optical illusion known as "motion dazzle," which makes it harder for predators to target a specific individual during a chase.\n\nFurthermore, being alone can lead to increased stress and a lack of opportunities for social interaction, grooming, and reproduction, which are vital for the overall well-being and survival of zebras. Consequently, a lone zebra in the wild may have a harder time thriving and ensuring the continuation of its lineage.</s>'], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000561594.jpg', tensor([[[-1.3473, -1.2617, -1.1075,  ..., -0.9534, -0.9705, -1.0048],
         [-1.0733, -0.9534, -0.7822,  ..., -0.9192, -0.9020, -0.9020],
         [-0.6965, -0.5424, -0.3369,  ..., -0.8678, -0.8164, -0.7650],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.3200, -0.4601, -0.6176,  ..., -0.9328, -0.8452, -0.7752],
         [-0.3725, -0.3550, -0.3375,  ..., -0.8803, -0.7752, -0.7052],
         [-0.3725, -0.1625,  0.1001,  ..., -0.7927, -0.6702, -0.6001],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.3164, -1.3339, -1.2816,  ..., -0.5147, -0.4624, -0.4450],
         [-1.0027, -0.8110, -0.4973,  ..., -0.4973, -0.4101, -0.3578],
         [-0.5670, -0.0615,  0.6182,  ..., -0.4450, -0.3230, -0.2532],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.8726, -0.9456, -0.5222,  ..., -1.1791, -1.3397, -1.3981],
         [-0.9456, -0.8872, -0.9748,  ..., -1.1791, -1.2959, -1.3689],
         [-1.1207, -0.8434, -1.0623,  ..., -1.1791, -1.1937, -1.2521],
         ...,
         [-1.6171, -1.6317, -1.6317,  ..., -0.8288, -1.0915, -0.7850],
         [-1.6171, -1.6171, -1.6171,  ..., -1.1061, -0.9164, -0.8142],
         [-1.6171, -1.6025, -1.6025,  ..., -1.0769, -0.8288, -0.8288]],

        [[-0.3264, -0.5065, -0.1313,  ..., -1.1218, -1.2568, -1.3019],
         [-0.4314, -0.2213, -0.4614,  ..., -1.1368, -1.2268, -1.2718],
         [-0.6865, -0.3114, -0.3864,  ..., -1.1218, -1.1368, -1.1968],
         ...,
         [-1.5570, -1.5720, -1.5720,  ..., -0.4014, -0.5815, -0.0862],
         [-1.5720, -1.5720, -1.5720,  ..., -0.6865, -0.3864, -0.2363],
         [-1.5570, -1.5570, -1.5570,  ..., -0.6415, -0.3114, -0.3714]],

        [[-0.4137, -0.4422, -0.1578,  ..., -0.8119, -0.8830, -0.8830],
         [-0.2289,  0.0413, -0.1720,  ..., -0.8403, -0.9114, -0.9683],
         [-0.5133,  0.0982, -0.1009,  ..., -0.8545, -0.8830, -0.8972],
         ...,
         [-1.3807, -1.3807, -1.3807,  ...,  0.3542,  0.0982,  0.6528],
         [-1.3807, -1.3807, -1.3807,  ...,  0.0413,  0.2262,  0.4395],
         [-1.3665, -1.3665, -1.3807,  ...,  0.1693,  0.3542,  0.3257]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the brown object next to the boat in this image. ASSISTANT: Result: brown object next to the boat.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the the wooden part of the boat where the woman with the blue shirt is touching in this image. ASSISTANT: The region corresponding to the the wooden part of the boat where the woman with the blue shirt is touching is shown.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the large canoe boat with bananas on top of green material from the provided image. ASSISTANT: Here is the large canoe boat with bananas on top of green material you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the boat full of banannas in this image? Please elaborate your answer and explain why. ASSISTANT: You can see the the boat full of banannas in this frame.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (683, 1024), ['<image>\nPlease segment the brown object next to the boat in this image.', '<image>\nPlease segment the the wooden part of the boat where the woman with the blue shirt is touching in this image.', '<image>\nSegment the large canoe boat with bananas on top of green material from the provided image.', '<image>\nWhat is the boat full of banannas in this image? Please elaborate your answer and explain why.'], ['brown object next to the boat', 'the wooden part of the boat where the woman with the blue shirt is touching', 'large canoe boat with bananas on top of green material', 'the boat full of banannas'], [{'brown object next to the boat': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the wooden part of the boat where the woman with the blue shirt is touching': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'large canoe boat with bananas on top of green material': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the boat full of banannas': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg')]
>> len(batch):  6
>> batch:  >> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000269504.jpg', tensor([[[ 1.9920,  1.9920,  1.9920,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.9920,  1.9920,  1.9920,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.9920,  1.9920,  1.9920,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 0.3138,  0.4166,  0.5878,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.3994,  0.4337,  0.4508,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.4508,  0.4337,  0.3823,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.4286,  2.4286,  2.4286,  ...,  0.0000,  0.0000,  0.0000],
         [ 2.4286,  2.4286,  2.4286,  ...,  0.0000,  0.0000,  0.0000],
         [ 2.4286,  2.4286,  2.4286,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 1.0630,  1.1506,  1.3782,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.1331,  1.1681,  1.2381,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.1681,  1.1681,  1.1681,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.6400,  2.6400,  2.6400,  ...,  0.0000,  0.0000,  0.0000],
         [ 2.6400,  2.6400,  2.6400,  ...,  0.0000,  0.0000,  0.0000],
         [ 2.6400,  2.6400,  2.6400,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.3230, -0.2532, -0.1312,  ...,  0.0000,  0.0000,  0.0000],
         [-0.2358, -0.2358, -0.2707,  ...,  0.0000,  0.0000,  0.0000],
         [-0.2010, -0.2358, -0.3230,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.9303,  1.9303,  1.9303,  ...,  1.8427,  1.8135,  1.7990],
         [ 1.9303,  1.9303,  1.9303,  ...,  1.9011,  1.8865,  1.8573],
         [ 1.9303,  1.9303,  1.9303,  ...,  1.9157,  1.9011,  1.8719],
         ...,
         [-0.3762, -0.4200, -0.1280,  ..., -0.3616, -0.2448, -0.0988],
         [-0.8580, -0.5222, -0.4638,  ..., -0.4638, -0.6974, -0.0259],
         [-0.8142, -0.6682, -0.4200,  ..., -0.6098, -0.6536, -0.3324]],

        [[ 2.0749,  2.0749,  2.0749,  ...,  2.0749,  2.0749,  2.0749],
         [ 2.0749,  2.0749,  2.0749,  ...,  2.0599,  2.0599,  2.0749],
         [ 2.0749,  2.0749,  2.0749,  ...,  2.0599,  2.0599,  2.0749],
         ...,
         [ 0.1839,  0.1239,  0.4390,  ...,  0.3490,  0.3940,  0.6191],
         [-0.2213,  0.0939,  0.0789,  ...,  0.3340, -0.0412,  0.5741],
         [-0.1613, -0.0262,  0.1989,  ...,  0.2289,  0.0488,  0.1839]],

        [[ 2.1459,  2.1459,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         [ 2.1459,  2.1459,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         [ 2.1459,  2.1459,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         ...,
         [-0.3995, -0.6128, -0.4706,  ..., -0.1009, -0.1862, -0.0867],
         [-0.7408, -0.5133, -0.5844,  ..., -0.1151, -0.5133,  0.0271],
         [-0.6697, -0.5417, -0.3284,  ..., -0.2715, -0.3853, -0.2573]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the fully visible giraffe in this image? ASSISTANT: Take a look at the fully visible giraffe here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the taller in this image. ASSISTANT: This is the selected taller.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is giraffe thats darker in this image? Please elaborate your answer and explain why. ASSISTANT: The giraffe thats darker is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the giraffe that is nearest tree and darker in this picture? ASSISTANT: giraffe that is nearest tree and darker.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the animal on side with most sky showing from the provided image. ASSISTANT: Here's where the animal on side with most sky showing appears in the image.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 768), ['<image>\nCan you segment the fully visible giraffe in this image?', '<image>\nPlease highlight the taller in this image.', '<image>\nWhat is giraffe thats darker in this image? Please elaborate your answer and explain why.', '<image>\nCould you identify and segment out the giraffe that is nearest tree and darker in this picture?', '<image>\nSegment the animal on side with most sky showing from the provided image.'], ['fully visible giraffe', 'taller', 'giraffe thats darker', 'giraffe that is nearest tree and darker', 'animal on side with most sky showing'], [{'fully visible giraffe': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'taller': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'giraffe thats darker': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'giraffe that is nearest tree and darker': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'animal on side with most sky showing': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000137420.jpg', tensor([[[ 1.0502,  1.0159,  0.9817,  ...,  0.3138,  0.3481,  0.3652],
         [ 1.0844,  1.0673,  1.0331,  ...,  0.2967,  0.3309,  0.3309],
         [ 1.1358,  1.1187,  1.1015,  ...,  0.2796,  0.2967,  0.2967],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.4328,  0.3978,  0.3627,  ..., -1.1954, -1.1779, -1.1604],
         [ 0.4678,  0.4503,  0.4153,  ..., -1.1954, -1.1779, -1.1779],
         [ 0.5203,  0.5028,  0.4853,  ..., -1.1954, -1.1954, -1.1954],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.0082, -0.0267, -0.0615,  ..., -1.6302, -1.6650, -1.6999],
         [ 0.0256,  0.0082, -0.0092,  ..., -1.6302, -1.6650, -1.7173],
         [ 0.0605,  0.0605,  0.0605,  ..., -1.6127, -1.6824, -1.7347],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.6457,  0.5289,  0.4559,  ...,  0.9230,  0.9230,  0.8938],
         [ 0.2515,  0.1785,  0.2369,  ...,  0.9230,  0.9084,  0.9084],
         [ 0.2369,  0.2661,  0.4851,  ...,  0.9084,  0.8938,  0.9084],
         ...,
         [ 0.8501,  0.8209,  0.8501,  ...,  0.2661,  0.1347,  0.1201],
         [ 0.2077,  0.1785,  0.1931,  ...,  0.5581,  0.5143,  0.5581],
         [-0.2448, -0.3032, -0.3178,  ...,  0.6457,  0.6749,  0.6603]],

        [[ 0.4841,  0.4090,  0.4240,  ..., -0.6265, -0.6265, -0.6415],
         [ 0.0789,  0.1539,  0.3190,  ..., -0.6265, -0.6415, -0.6415],
         [ 0.0789,  0.3490,  0.7542,  ..., -0.6265, -0.6415, -0.6415],
         ...,
         [ 0.7392,  0.7392,  0.7842,  ..., -0.2663, -0.4764, -0.5215],
         [ 0.0188,  0.0188,  0.0488,  ...,  0.0638, -0.0112,  0.0188],
         [-0.3264, -0.4314, -0.4164,  ...,  0.1839,  0.1389,  0.1539]],

        [[ 0.3542,  0.5248,  0.6528,  ..., -1.1816, -1.1958, -1.2385],
         [ 0.3542,  0.7950,  1.1505,  ..., -1.1958, -1.2100, -1.2100],
         [ 0.6386,  1.3069,  1.7904,  ..., -1.2811, -1.2527, -1.1816],
         ...,
         [ 0.9514,  0.9941,  1.0083,  ..., -0.4848, -0.6839, -0.9541],
         [ 0.4253,  0.4821,  0.5106,  ..., -0.2715, -0.3284, -0.3853],
         [-0.3000, -0.3284, -0.2857,  ..., -0.1435, -0.0724, -0.1151]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat type of bird is in the image? ASSISTANT: The image features a colorful green parrot, which has blue and red markings on its wings.</s>USER: Where is the parrot located in relation to the cage? ASSISTANT: The parrot is standing on top of the cage, outside of it.</s>USER: Is the parrot posing for a close-up shot or far away from the camera? ASSISTANT: The parrot is posing for a close-up shot. It is looking directly into the camera, which captures the vibrant colors and intricate details of its feathers.</s>USER: Where is the cage with the parrot located? ASSISTANT: The cage with the parrot on top is located inside a residence or a home.</s>USER: What colors are visible on the parrot's wings? ASSISTANT: Blue and red colors can be seen on the parrot's wings, in addition to its bright green body.</s>"], tensor([], size=(0, 481, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (770, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat type of bird is in the image? ASSISTANT: The image features a colorful green parrot, which has blue and red markings on its wings.</s>USER: Where is the parrot located in relation to the cage? ASSISTANT: The parrot is standing on top of the cage, outside of it.</s>USER: Is the parrot posing for a close-up shot or far away from the camera? ASSISTANT: The parrot is posing for a close-up shot. It is looking directly into the camera, which captures the vibrant colors and intricate details of its feathers.</s>USER: Where is the cage with the parrot located? ASSISTANT: The cage with the parrot on top is located inside a residence or a home.</s>USER: What colors are visible on the parrot's wings? ASSISTANT: Blue and red colors can be seen on the parrot's wings, in addition to its bright green body.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat type of bird is in the image? ASSISTANT: The image features a colorful green parrot, which has blue and red markings on its wings.</s>USER: Where is the parrot located in relation to the cage? ASSISTANT: The parrot is standing on top of the cage, outside of it.</s>USER: Is the parrot posing for a close-up shot or far away from the camera? ASSISTANT: The parrot is posing for a close-up shot. It is looking directly into the camera, which captures the vibrant colors and intricate details of its feathers.</s>USER: Where is the cage with the parrot located? ASSISTANT: The cage with the parrot on top is located inside a residence or a home.</s>USER: What colors are visible on the parrot's wings? ASSISTANT: Blue and red colors can be seen on the parrot's wings, in addition to its bright green body.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2009_000367.jpg', tensor([[[-0.0116, -0.0116, -0.0116,  ..., -0.0801, -0.0629, -0.0629],
         [-0.0116, -0.0116, -0.0116,  ..., -0.0801, -0.0629, -0.0629],
         [ 0.0056,  0.0056,  0.0056,  ..., -0.0629, -0.0458, -0.0458],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.2752,  0.2752,  0.2927,  ...,  0.2402,  0.2577,  0.2577],
         [ 0.2927,  0.2927,  0.2927,  ...,  0.2402,  0.2577,  0.2577],
         [ 0.3102,  0.3102,  0.3102,  ...,  0.2577,  0.2752,  0.2752],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.0017,  0.9668,  0.8971,  ...,  0.7751,  0.7925,  0.7925],
         [ 1.0017,  0.9668,  0.8971,  ...,  0.7751,  0.7925,  0.7925],
         [ 0.9842,  0.9494,  0.8971,  ...,  0.7925,  0.8099,  0.8099],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.0471,  0.0471,  0.0763,  ..., -0.0113,  0.0033,  0.0471],
         [ 0.0471,  0.0179,  0.0179,  ...,  0.0179,  0.0325,  0.0033],
         [ 0.0325,  0.0179,  0.0179,  ...,  0.0179, -0.0113,  0.0179],
         ...,
         [-0.0988, -0.2302, -0.0405,  ...,  0.5727,  0.4413,  0.5143],
         [-0.0988, -0.3470, -0.2010,  ...,  0.4267,  0.3537,  0.3537],
         [ 0.2223,  0.2807,  0.2369,  ...,  0.1347,  0.0909,  0.1639]],

        [[ 0.2890,  0.3040,  0.3340,  ...,  0.2890,  0.2589,  0.2589],
         [ 0.2890,  0.3190,  0.3340,  ...,  0.2890,  0.2890,  0.3040],
         [ 0.2890,  0.3340,  0.3340,  ...,  0.2740,  0.2740,  0.2589],
         ...,
         [-0.3264, -0.4164, -0.3414,  ..., -0.1313, -0.2813, -0.2063],
         [-0.4914, -0.5065, -0.4314,  ..., -0.1913, -0.2813, -0.3564],
         [-0.4314, -0.3264, -0.3714,  ..., -0.4014, -0.4164, -0.4014]],

        [[ 0.7381,  0.7523,  0.7381,  ...,  0.6670,  0.6670,  0.6812],
         [ 0.7381,  0.7239,  0.7097,  ...,  0.6528,  0.6670,  0.6670],
         [ 0.7523,  0.7381,  0.7381,  ...,  0.6528,  0.6528,  0.6528],
         ...,
         [-0.7692, -0.8403, -0.7266,  ..., -0.6412, -0.7692, -0.8403],
         [-0.8119, -0.8972, -0.7550,  ..., -0.6981, -0.7266, -0.7834],
         [-0.7550, -0.7834, -0.7834,  ..., -0.8261, -0.8830, -0.8261]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the aeroplane wheel shown in this image. ASSISTANT: Below you can see the aeroplane wheel.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the aeroplane wing in this picture? ASSISTANT: Take a look at the aeroplane wing here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the aeroplane tail from the provided image. ASSISTANT: You can see the aeroplane tail in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the the body of the aeroplane from the image below? ASSISTANT: Below you can see the the body of the aeroplane.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the wheel of the aeroplane in this image? Please elaborate your answer and explain why. ASSISTANT: I've marked the the wheel of the aeroplane for you.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (666, 1024), ['<image>\nDisplay a segmentation mask for the aeroplane wheel shown in this image.', '<image>\nCould you identify and segment out the aeroplane wing in this picture?', '<image>\nSegment the aeroplane tail from the provided image.', '<image>\nWould you please extract the the body of the aeroplane from the image below?', '<image>\nWhat is the wheel of the aeroplane in this image? Please elaborate your answer and explain why.'], ['aeroplane wheel', 'aeroplane wing', 'aeroplane tail', 'the body of the aeroplane', 'the wheel of the aeroplane'], [{'aeroplane wheel': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'aeroplane wing': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'aeroplane tail': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the body of the aeroplane': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the wheel of the aeroplane': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000393985.jpg', tensor([[[-1.8268, -1.7069, -1.5528,  ..., -1.8268, -1.8268, -1.8268],
         [-1.8268, -1.7069, -1.5528,  ..., -1.8097, -1.8268, -1.8268],
         [-1.8097, -1.7069, -1.5528,  ..., -1.7925, -1.8097, -1.8097],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.7556, -1.6681, -1.5280,  ..., -1.6331, -1.6331, -1.6331],
         [-1.7381, -1.6681, -1.5280,  ..., -1.6155, -1.6331, -1.6331],
         [-1.7206, -1.6506, -1.5280,  ..., -1.5980, -1.6155, -1.6155],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.5604, -1.4733, -1.3513,  ..., -1.4384, -1.4384, -1.4384],
         [-1.5256, -1.4559, -1.3513,  ..., -1.4210, -1.4384, -1.4384],
         [-1.4907, -1.4384, -1.3513,  ..., -1.4036, -1.4210, -1.4210],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.1353e+00, -1.1499e+00, -8.2877e-01,  ..., -1.0623e+00,
          -1.0623e+00, -1.1207e+00],
         [-1.3397e+00, -1.3251e+00, -1.4273e+00,  ..., -1.2083e+00,
          -1.1791e+00, -1.1499e+00],
         [-1.3397e+00, -1.1353e+00, -1.1791e+00,  ..., -1.0769e+00,
          -1.1207e+00, -1.1499e+00],
         ...,
         [ 4.2670e-01, -2.5853e-02,  3.5371e-01,  ..., -5.9519e-01,
          -4.3461e-01, -5.5140e-01],
         [-1.7184e-01, -2.0103e-01,  3.2541e-02,  ..., -4.9300e-01,
          -2.0103e-01, -6.0979e-01],
         [ 1.7853e-01,  3.2541e-02,  4.7139e-02,  ..., -6.9738e-01,
          -5.9519e-01, -6.2439e-01]],

        [[-1.0317e+00, -9.8670e-01, -5.8149e-01,  ..., -9.8670e-01,
          -1.0467e+00, -1.1068e+00],
         [-1.1368e+00, -1.2268e+00, -1.2869e+00,  ..., -1.2718e+00,
          -1.2568e+00, -1.1368e+00],
         [-1.1968e+00, -9.2667e-01, -1.0167e+00,  ..., -1.0617e+00,
          -1.1218e+00, -1.0918e+00],
         ...,
         [ 5.4409e-01,  7.8851e-02,  4.6905e-01,  ..., -4.6143e-01,
          -3.2636e-01, -3.2636e-01],
         [-7.1227e-02, -1.1625e-01,  1.3888e-01,  ..., -3.4137e-01,
          -7.1227e-02, -4.1641e-01],
         [ 2.7395e-01,  1.3888e-01,  1.6890e-01,  ..., -5.3647e-01,
          -4.4642e-01, -4.3141e-01]],

        [[-8.2610e-01, -7.6922e-01, -4.2793e-01,  ..., -8.6876e-01,
          -8.9720e-01, -9.2564e-01],
         [-8.9720e-01, -9.2564e-01, -1.0252e+00,  ..., -1.0536e+00,
          -1.0536e+00, -9.6830e-01],
         [-9.8252e-01, -6.8390e-01, -7.6922e-01,  ..., -8.8298e-01,
          -9.3986e-01, -9.3986e-01],
         ...,
         [ 5.6747e-01,  1.5509e-01,  5.3903e-01,  ..., -3.7105e-01,
          -2.5729e-01, -2.4307e-01],
         [ 4.1327e-02, -1.3329e-03,  1.9775e-01,  ..., -2.5729e-01,
          -1.3329e-03, -3.4261e-01],
         [ 3.6839e-01,  2.2619e-01,  2.1197e-01,  ..., -4.2793e-01,
          -3.4261e-01, -3.4261e-01]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nAre the people walking in the rain using any protection from the wet weather? ASSISTANT: Yes, the people walking in the rain are using umbrellas to protect themselves from the wet weather.</s>USER: How many people are in the group walking across the street? ASSISTANT: Although there is a group of people walking across the street, I cannot provide a specific number of individuals in the group without an exact count from the image.</s>USER: Are the people walking on a sidewalk, or are they crossing the street? ASSISTANT: The people are walking across the street, while using their umbrellas to shield themselves from the rain.</s>USER: What kind of environment is this scene happening in? Urban, suburban, or rural setting? ASSISTANT: Based on the details provided, the scene is most likely happening in an urban or suburban setting. This would be a reasonable assumption considering that there is a street, a group of people walking together with umbrellas, and the fact that they are crossing the street, which are features commonly found in urban or suburban settings rather than rural ones.</s>USER: What are some safety tips for pedestrians walking in the rain? ASSISTANT: Here are a few safety tips for pedestrians walking in the rain:\n\n1. Use appropriate rain gear: Carry an umbrella or wear a waterproof jacket to stay dry in the rain. Wearing appropriate footwear, such as rain boots or shoes with a good grip, can help prevent slipping on wet surfaces.\n\n2. Increase visibility: Wear brightly colored clothing, reflective gear, or use accessories like lights or reflective bands to increase visibility for drivers and other pedestrians, as rain and wet conditions can impair visibility.\n\n3. Look before crossing: When crossing the street, look both ways and make eye contact with drivers, if possible, to ensure they see you before you proceed.\n\n4. Be cautious on slippery surfaces: Slow down and take extra care when walking on wet, slippery surfaces or through puddles, as slipping and falling can cause injuries.\n\n5. Follow traffic rules: Obey traffic signals, cross at designated crosswalks or intersections, and be mindful of turning vehicles.\n\n6. Stay alert: Keep your ears open and remain aware of your surroundings whilst walking in the rain. Avoid distractions like using mobile devices or wearing headphones, which can impair your ability to hear oncoming traffic or potential hazards.\n\nBy following these safety tips, pedestrians can reduce the risks associated with walking in the rain and ensure a safer, more comfortable experience.</s>"], tensor([], size=(0, 427, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (683, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nAre the people walking in the rain using any protection from the wet weather? ASSISTANT: Yes, the people walking in the rain are using umbrellas to protect themselves from the wet weather.</s>USER: How many people are in the group walking across the street? ASSISTANT: Although there is a group of people walking across the street, I cannot provide a specific number of individuals in the group without an exact count from the image.</s>USER: Are the people walking on a sidewalk, or are they crossing the street? ASSISTANT: The people are walking across the street, while using their umbrellas to shield themselves from the rain.</s>USER: What kind of environment is this scene happening in? Urban, suburban, or rural setting? ASSISTANT: Based on the details provided, the scene is most likely happening in an urban or suburban setting. This would be a reasonable assumption considering that there is a street, a group of people walking together with umbrellas, and the fact that they are crossing the street, which are features commonly found in urban or suburban settings rather than rural ones.</s>USER: What are some safety tips for pedestrians walking in the rain? ASSISTANT: Here are a few safety tips for pedestrians walking in the rain:\n\n1. Use appropriate rain gear: Carry an umbrella or wear a waterproof jacket to stay dry in the rain. Wearing appropriate footwear, such as rain boots or shoes with a good grip, can help prevent slipping on wet surfaces.\n\n2. Increase visibility: Wear brightly colored clothing, reflective gear, or use accessories like lights or reflective bands to increase visibility for drivers and other pedestrians, as rain and wet conditions can impair visibility.\n\n3. Look before crossing: When crossing the street, look both ways and make eye contact with drivers, if possible, to ensure they see you before you proceed.\n\n4. Be cautious on slippery surfaces: Slow down and take extra care when walking on wet, slippery surfaces or through puddles, as slipping and falling can cause injuries.\n\n5. Follow traffic rules: Obey traffic signals, cross at designated crosswalks or intersections, and be mindful of turning vehicles.\n\n6. Stay alert: Keep your ears open and remain aware of your surroundings whilst walking in the rain. Avoid distractions like using mobile devices or wearing headphones, which can impair your ability to hear oncoming traffic or potential hazards.\n\nBy following these safety tips, pedestrians can reduce the risks associated with walking in the rain and ensure a safer, more comfortable experience.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nAre the people walking in the rain using any protection from the wet weather? ASSISTANT: Yes, the people walking in the rain are using umbrellas to protect themselves from the wet weather.</s>USER: How many people are in the group walking across the street? ASSISTANT: Although there is a group of people walking across the street, I cannot provide a specific number of individuals in the group without an exact count from the image.</s>USER: Are the people walking on a sidewalk, or are they crossing the street? ASSISTANT: The people are walking across the street, while using their umbrellas to shield themselves from the rain.</s>USER: What kind of environment is this scene happening in? Urban, suburban, or rural setting? ASSISTANT: Based on the details provided, the scene is most likely happening in an urban or suburban setting. This would be a reasonable assumption considering that there is a street, a group of people walking together with umbrellas, and the fact that they are crossing the street, which are features commonly found in urban or suburban settings rather than rural ones.</s>USER: What are some safety tips for pedestrians walking in the rain? ASSISTANT: Here are a few safety tips for pedestrians walking in the rain:\n\n1. Use appropriate rain gear: Carry an umbrella or wear a waterproof jacket to stay dry in the rain. Wearing appropriate footwear, such as rain boots or shoes with a good grip, can help prevent slipping on wet surfaces.\n\n2. Increase visibility: Wear brightly colored clothing, reflective gear, or use accessories like lights or reflective bands to increase visibility for drivers and other pedestrians, as rain and wet conditions can impair visibility.\n\n3. Look before crossing: When crossing the street, look both ways and make eye contact with drivers, if possible, to ensure they see you before you proceed.\n\n4. Be cautious on slippery surfaces: Slow down and take extra care when walking on wet, slippery surfaces or through puddles, as slipping and falling can cause injuries.\n\n5. Follow traffic rules: Obey traffic signals, cross at designated crosswalks or intersections, and be mindful of turning vehicles.\n\n6. Stay alert: Keep your ears open and remain aware of your surroundings whilst walking in the rain. Avoid distractions like using mobile devices or wearing headphones, which can impair your ability to hear oncoming traffic or potential hazards.\n\nBy following these safety tips, pedestrians can reduce the risks associated with walking in the rain and ensure a safer, more comfortable experience.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000178821.jpg', tensor([[[-1.1932, -1.0562, -0.8507,  ...,  0.0000,  0.0000,  0.0000],
         [-1.1932, -1.0733, -0.8678,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2103, -1.0733, -0.8849,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.6384, -1.6213, -1.5870,  ...,  0.0000,  0.0000,  0.0000],
         [-1.5699, -1.5528, -1.5185,  ...,  0.0000,  0.0000,  0.0000],
         [-1.5185, -1.5014, -1.4672,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.2654, -1.2654, -1.2654,  ...,  0.0000,  0.0000,  0.0000],
         [-1.3004, -1.3179, -1.3179,  ...,  0.0000,  0.0000,  0.0000],
         [-1.3529, -1.3704, -1.3704,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.6331, -1.6506, -1.6681,  ...,  0.0000,  0.0000,  0.0000],
         [-1.5630, -1.5805, -1.5980,  ...,  0.0000,  0.0000,  0.0000],
         [-1.4930, -1.5105, -1.5280,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.2467, -1.2641, -1.2816,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2816, -1.2990, -1.3164,  ...,  0.0000,  0.0000,  0.0000],
         [-1.3164, -1.3513, -1.3513,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.4907, -1.4907, -1.4907,  ...,  0.0000,  0.0000,  0.0000],
         [-1.4733, -1.4559, -1.4559,  ...,  0.0000,  0.0000,  0.0000],
         [-1.4559, -1.4384, -1.4210,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.0033, -0.4638, -1.6171,  ...,  0.9960,  0.9960,  1.0106],
         [ 0.0033, -0.5076, -1.5879,  ...,  0.9668,  0.9814,  0.9814],
         [-0.0113, -0.5660, -1.5879,  ...,  0.9376,  0.9522,  0.9668],
         ...,
         [-0.5660, -0.5806, -0.5514,  ..., -1.4565, -1.4711, -1.5149],
         [-0.6974, -0.6682, -0.6098,  ..., -1.4419, -1.4711, -1.5003],
         [-0.6390, -0.5806, -0.5368,  ..., -1.5149, -1.5149, -1.4711]],

        [[-0.7466, -0.9867, -1.6170,  ...,  0.9643,  0.9643,  0.9793],
         [-0.7016, -1.0317, -1.6170,  ...,  0.9493,  0.9643,  0.9643],
         [-0.7166, -1.0767, -1.6170,  ...,  0.9043,  0.9343,  0.9493],
         ...,
         [-1.1068, -1.0918, -1.0467,  ..., -1.4970, -1.4820, -1.4970],
         [-1.1518, -1.1218, -1.0918,  ..., -1.5270, -1.4970, -1.4669],
         [-1.1518, -1.1368, -1.1068,  ..., -1.5120, -1.4970, -1.4970]],

        [[-0.9683, -1.1247, -1.3807,  ...,  1.0652,  1.0510,  1.0652],
         [-0.9825, -1.1389, -1.3807,  ...,  1.0367,  1.0510,  1.0510],
         [-0.9399, -1.1816, -1.3949,  ...,  1.0225,  1.0225,  1.0083],
         ...,
         [-1.2385, -1.2385, -1.1958,  ..., -1.2243, -1.1958, -1.2385],
         [-1.2385, -1.2527, -1.2243,  ..., -1.2527, -1.2243, -1.2100],
         [-1.1389, -1.1958, -1.2100,  ..., -1.2527, -1.2385, -1.1958]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is right man in this image? Please give some explanation. ASSISTANT: Here is the right man you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the man front in this picture? ASSISTANT: The man front is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the right guy from the image below? ASSISTANT: This is the selected right guy.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the front person in this image. ASSISTANT: You can see the front person in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the guy far right in this picture? ASSISTANT: This is the selected guy far right.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 722), ['<image>\nWhat is right man in this image? Please give some explanation.', '<image>\nCould you isolate the man front in this picture?', '<image>\nWould you please extract the right guy from the image below?', '<image>\nPlease highlight the front person in this image.', '<image>\nCould you isolate the guy far right in this picture?'], ['right man', 'man front', 'right guy', 'front person', 'guy far right'], [{'right man': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'man front': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'right guy': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'front person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'guy far right': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/saiapr_tc-12/01/images/1357.jpg', tensor([[[ 2.0434,  2.0605,  2.0777,  ..., -1.4672, -1.4672, -1.4672],
         [ 2.0605,  2.0605,  2.0777,  ..., -1.4672, -1.4672, -1.4672],
         [ 2.0777,  2.0777,  2.0777,  ..., -1.4672, -1.4672, -1.4672],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.4111,  2.4111,  2.4286,  ..., -1.7031, -1.7031, -1.7031],
         [ 2.4111,  2.4111,  2.4286,  ..., -1.7031, -1.7031, -1.7031],
         [ 2.4286,  2.4286,  2.4286,  ..., -1.7031, -1.7031, -1.7031],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.5529,  2.5703,  2.5877,  ..., -1.5779, -1.5779, -1.5779],
         [ 2.5703,  2.5877,  2.6051,  ..., -1.5779, -1.5779, -1.5779],
         [ 2.6051,  2.6226,  2.6226,  ..., -1.5779, -1.5779, -1.5779],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.3610,  1.2734,  1.2004,  ..., -0.7996, -0.8434, -0.8434],
         [ 1.4048,  1.3026,  1.2442,  ..., -0.7412, -0.7996, -0.7996],
         [ 1.4486,  1.3318,  1.2880,  ..., -0.7558, -0.8288, -0.8288],
         ...,
         [ 1.8573,  1.8427,  1.8281,  ..., -0.7558, -0.7120, -0.7704],
         [ 1.8719,  1.8427,  1.7990,  ..., -0.7850, -0.7412, -0.7704],
         [ 1.7990,  1.7552,  1.7114,  ..., -0.7558, -0.6974, -0.7558]],

        [[ 1.3845,  1.3695,  1.3695,  ..., -1.4219, -1.4219, -1.4369],
         [ 1.4596,  1.4145,  1.3995,  ..., -1.3619, -1.3919, -1.3919],
         [ 1.4896,  1.4446,  1.4446,  ..., -1.4069, -1.4369, -1.4369],
         ...,
         [ 1.9848,  1.9548,  1.9098,  ..., -0.8366, -0.8066, -0.8516],
         [ 1.9998,  1.9548,  1.8948,  ..., -0.8666, -0.8216, -0.8516],
         [ 1.9248,  1.8648,  1.8047,  ..., -0.8366, -0.7766, -0.8366]],

        [[ 1.1363,  1.1363,  1.0936,  ..., -1.2100, -1.2100, -1.2243],
         [ 1.2074,  1.1789,  1.1363,  ..., -1.1532, -1.1816, -1.1816],
         [ 1.2785,  1.2500,  1.1932,  ..., -1.1816, -1.2243, -1.2100],
         ...,
         [ 2.1459,  2.1459,  2.1459,  ..., -0.5133, -0.4564, -0.4706],
         [ 2.1459,  2.1459,  2.1317,  ..., -0.5417, -0.4848, -0.4706],
         [ 2.1175,  2.1175,  2.0890,  ..., -0.4990, -0.4422, -0.4422]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is brown curtain near lamp and green wall in this image? Please elaborate your answer and explain why. ASSISTANT: This view focuses on the brown curtain near lamp and green wall.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the green wall on left from the provided image. ASSISTANT: Here's where the green wall on left appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the the dark door in this picture? ASSISTANT: Here's the area for the dark door.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the lamp from the provided image. ASSISTANT: Here is the lamp you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the orange wall behind bed shown in this image. ASSISTANT: Below you can see the orange wall behind bed.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 1, 1, 1],
         [0, 0, 0,  ..., 1, 1, 1],
         [0, 0, 0,  ..., 1, 1, 1],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nWhat is brown curtain near lamp and green wall in this image? Please elaborate your answer and explain why.', '<image>\nSegment the green wall on left from the provided image.', '<image>\nCould you isolate the the dark door in this picture?', '<image>\nSegment the lamp from the provided image.', '<image>\nDisplay a segmentation mask for the orange wall behind bed shown in this image.'], ['brown curtain near lamp and green wall', 'green wall on left', 'the dark door', 'lamp', 'orange wall behind bed'], [{'brown curtain near lamp and green wall': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'green wall on left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the dark door': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'lamp': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'orange wall behind bed': tensor([[0, 0, 0,  ..., 1, 1, 1],
        [0, 0, 0,  ..., 1, 1, 1],
        [0, 0, 0,  ..., 1, 1, 1],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg')]
[('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2008_002181.jpg', tensor([[[-0.5767, -0.5767, -0.5938,  ..., -1.8268, -1.8268, -1.8268],
         [-0.7650, -0.7650, -0.7822,  ..., -1.8268, -1.8268, -1.8268],
         [-1.1589, -1.1589, -1.1760,  ..., -1.8268, -1.8268, -1.8268],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.4601, -0.4776, -0.4951,  ..., -1.6681, -1.6681, -1.6681],
         [-0.6527, -0.6702, -0.6877,  ..., -1.6681, -1.6681, -1.6681],
         [-1.0728, -1.0903, -1.1078,  ..., -1.6681, -1.6681, -1.6681],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.6541, -0.6541, -0.6541,  ..., -1.3861, -1.3861, -1.3861],
         [-0.8110, -0.8110, -0.8110,  ..., -1.3861, -1.3861, -1.3861],
         [-1.1596, -1.1596, -1.1596,  ..., -1.3861, -1.3861, -1.3861],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.1499, -1.1499, -1.1645,  ..., -1.5149, -1.5441, -1.5441],
         [-1.1499, -1.1499, -1.1207,  ..., -1.5149, -1.5441, -1.5441],
         [-1.1499, -1.1499, -1.1061,  ..., -1.5441, -1.5441, -1.5441],
         ...,
         [ 0.1201,  0.0325,  0.1201,  ..., -1.4711, -1.4419, -1.4565],
         [ 0.2515,  0.1639,  0.1931,  ..., -1.4565, -1.4273, -1.4565],
         [ 0.1931,  0.2223,  0.2515,  ..., -1.4711, -1.4127, -1.4419]],

        [[-1.1518, -1.1518, -1.1668,  ..., -1.4069, -1.4369, -1.4369],
         [-1.1518, -1.1518, -1.1218,  ..., -1.4219, -1.4369, -1.4369],
         [-1.1518, -1.1518, -1.1068,  ..., -1.4219, -1.4369, -1.4369],
         ...,
         [ 0.1839,  0.0789,  0.1089,  ..., -1.3769, -1.3469, -1.3619],
         [ 0.3190,  0.2139,  0.1839,  ..., -1.3619, -1.3319, -1.3619],
         [ 0.2289,  0.2589,  0.2289,  ..., -1.3769, -1.3169, -1.3469]],

        [[-0.9541, -0.9541, -0.9683,  ..., -1.1247, -1.1389, -1.1389],
         [-0.9541, -0.9541, -0.9256,  ..., -1.1247, -1.1389, -1.1389],
         [-0.9541, -0.9541, -0.9114,  ..., -1.1247, -1.1389, -1.1389],
         ...,
         [ 0.0840,  0.0413, -0.0156,  ..., -1.0536, -1.0252, -1.0394],
         [ 0.1977,  0.0840,  0.0982,  ..., -1.0394, -1.0110, -1.0394],
         [ 0.1409,  0.1266,  0.1977,  ..., -1.0536, -0.9967, -1.0252]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the person leg in this picture? ASSISTANT: person leg.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the the leg of the person from the provided image. ASSISTANT: Sure — this shows the the leg of the person.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the person neck from the provided image. ASSISTANT: I've marked the person neck for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the person leg in this image? ASSISTANT: The person leg is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the the leg of the person from the provided image. ASSISTANT: Here is the the leg of the person you asked about.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (930, 1024), ['<image>\nCould you identify and segment out the person leg in this picture?', '<image>\nSegment the the leg of the person from the provided image.', '<image>\nSegment the person neck from the provided image.', '<image>\nCan you segment the person leg in this image?', '<image>\nSegment the the leg of the person from the provided image.'], ['person leg', 'the leg of the person', 'person neck', 'person leg', 'the leg of the person'], [{'person leg': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the leg of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person neck': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person leg': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the leg of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/YxDoMS_KCBHdxeO8p3igoA.jpg', tensor([[[1.5810, 1.5810, 1.6153,  ..., 0.8104, 0.7933, 0.7933],
         [1.5810, 1.5810, 1.5982,  ..., 0.7933, 0.7762, 0.7933],
         [1.5982, 1.5982, 1.5982,  ..., 0.7933, 0.7762, 0.7933],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.0609, 2.0609, 2.0959,  ..., 1.3081, 1.2906, 1.2906],
         [2.0609, 2.0609, 2.0784,  ..., 1.2906, 1.2731, 1.2906],
         [2.0784, 2.0784, 2.0784,  ..., 1.2906, 1.2731, 1.2906],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.3437, 2.3437, 2.3786,  ..., 1.6814, 1.6640, 1.6640],
         [2.3437, 2.3437, 2.3611,  ..., 1.6640, 1.6465, 1.6640],
         [2.3611, 2.3611, 2.3611,  ..., 1.6640, 1.6465, 1.6640],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 3.9750e-01,  2.9531e-01,  3.0991e-01,  ...,  9.3764e-01,
           9.3764e-01,  9.0845e-01],
         [ 4.4130e-01,  5.2889e-01,  5.5808e-01,  ...,  8.5005e-01,
           8.5005e-01,  8.2086e-01],
         [ 4.7049e-01,  5.5808e-01,  4.9969e-01,  ...,  7.7706e-01,
           7.6246e-01,  7.4786e-01],
         ...,
         [-2.1563e-01, -2.5943e-01, -2.0103e-01,  ..., -3.1782e-01,
          -3.4702e-01, -3.9081e-01],
         [-2.5943e-01, -2.8862e-01, -3.3242e-01,  ..., -4.2001e-01,
          -2.7403e-01, -3.4702e-01],
         [-1.7184e-01, -2.7403e-01, -3.9081e-01,  ..., -4.9300e-01,
          -2.7403e-01, -2.0103e-01]],

        [[ 9.9432e-01,  9.0428e-01,  9.0428e-01,  ...,  1.3095e+00,
           1.3095e+00,  1.3095e+00],
         [ 1.0393e+00,  1.1444e+00,  1.1744e+00,  ...,  1.2344e+00,
           1.2344e+00,  1.2194e+00],
         [ 1.0694e+00,  1.1744e+00,  1.1144e+00,  ...,  1.1744e+00,
           1.1594e+00,  1.1594e+00],
         ...,
         [-4.1212e-02, -7.1227e-02, -2.6204e-02,  ..., -1.4627e-01,
          -2.0630e-01, -2.5132e-01],
         [-8.6235e-02, -1.3126e-01, -1.6127e-01,  ..., -2.5132e-01,
          -1.3126e-01, -2.0630e-01],
         [-1.1196e-02, -1.1625e-01, -2.2130e-01,  ..., -3.2636e-01,
          -1.3126e-01, -5.6219e-02]],

        [[ 1.5629e+00,  1.4349e+00,  1.4776e+00,  ...,  1.5344e+00,
           1.5344e+00,  1.5202e+00],
         [ 1.6055e+00,  1.6624e+00,  1.6766e+00,  ...,  1.4633e+00,
           1.4491e+00,  1.4349e+00],
         [ 1.6198e+00,  1.7051e+00,  1.6340e+00,  ...,  1.3922e+00,
           1.3780e+00,  1.3780e+00],
         ...,
         [-1.3329e-03, -4.3993e-02,  1.2887e-02,  ..., -1.7197e-01,
          -2.2885e-01, -2.5729e-01],
         [-2.9773e-02, -5.8213e-02, -1.1509e-01,  ..., -2.5729e-01,
          -1.4353e-01, -2.1463e-01],
         [ 6.9767e-02, -2.9773e-02, -1.5775e-01,  ..., -3.2839e-01,
          -1.4353e-01, -7.2433e-02]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the pole in this picture? ASSISTANT: You can see the pole in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the sidewalk shown in this image. ASSISTANT: I've marked the sidewalk for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is sky in this image? Please give some explanation. ASSISTANT: I've marked the sky for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the road from the provided image. ASSISTANT: The road portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the mountain from the image below? ASSISTANT: Here's where the mountain appears in the image.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        ...,
        [21, 21, 21,  ..., 21, 21, 21],
        [21, 21, 21,  ..., 21, 21, 21],
        [21, 21, 21,  ..., 21, 21, 21]]), (768, 1024), ['<image>\nCould you identify and segment out the pole in this picture?', '<image>\nDisplay a segmentation mask for the sidewalk shown in this image.', '<image>\nWhat is sky in this image? Please give some explanation.', '<image>\nSegment the road from the provided image.', '<image>\nWould you please extract the mountain from the image below?'], ['pole', 'sidewalk', 'sky', 'road', 'mountain'], [{'pole': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'sidewalk': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'sky': tensor([[ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'road': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True]])}, {'mountain': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2008_002668.jpg', tensor([[[-0.6109, -0.6452, -0.7137,  ..., -1.2788, -1.2617, -1.2617],
         [-0.6452, -0.6794, -0.7479,  ..., -1.2788, -1.2617, -1.2617],
         [-0.7308, -0.7479, -0.8164,  ..., -1.2617, -1.2445, -1.2445],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.8277, -0.8627, -0.9153,  ..., -1.3880, -1.4055, -1.4055],
         [-0.8277, -0.8627, -0.9153,  ..., -1.3880, -1.4055, -1.4055],
         [-0.8452, -0.8803, -0.9328,  ..., -1.3880, -1.3880, -1.3880],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.8981, -0.9330, -1.0027,  ..., -1.3687, -1.3687, -1.3687],
         [-0.9156, -0.9504, -1.0201,  ..., -1.3687, -1.3687, -1.3687],
         [-0.9678, -0.9853, -1.0376,  ..., -1.3861, -1.3861, -1.3861],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.5952, -0.6244, -0.6098,  ..., -0.8288, -0.8434, -0.8580],
         [-0.5806, -0.5952, -0.6244,  ..., -0.8288, -0.8288, -0.8580],
         [-0.6098, -0.6390, -0.6828,  ..., -0.7996, -0.7558, -0.7704],
         ...,
         [ 1.7844,  1.8865,  1.8719,  ..., -0.7704, -0.8142, -0.8580],
         [ 1.4048,  1.8865,  1.8719,  ..., -0.8434, -0.8580, -0.9164],
         [ 1.3756,  1.9011,  1.8865,  ..., -0.8872, -0.9164, -0.9456]],

        [[-0.7616, -0.7916, -0.7916,  ..., -1.0317, -1.0617, -1.0167],
         [-0.7316, -0.7616, -0.7766,  ..., -1.0317, -1.0467, -1.0617],
         [-0.7616, -0.7916, -0.8066,  ..., -1.0017, -1.0017, -1.0167],
         ...,
         [ 1.1144,  1.4295,  1.6547,  ..., -1.0017, -1.0017, -1.0467],
         [ 0.6491,  1.3845,  1.7297,  ..., -1.0617, -1.0317, -1.1068],
         [ 0.5291,  1.3395,  1.5646,  ..., -1.0918, -1.1068, -1.1818]],

        [[-0.8688, -0.8972, -0.8972,  ..., -1.0394, -1.0678, -1.0536],
         [-0.8403, -0.8688, -0.8830,  ..., -1.0394, -1.0536, -1.0678],
         [-0.8688, -0.8972, -0.9256,  ..., -1.0536, -1.0110, -1.0110],
         ...,
         [ 0.3826,  0.7097,  0.9941,  ..., -0.9683, -0.9683, -1.0536],
         [-0.0156,  0.6812,  0.9941,  ..., -1.0252, -0.9825, -1.0678],
         [-0.1151,  0.6101,  0.8661,  ..., -1.0536, -0.9825, -1.0678]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the the nose of the cat in this image. ASSISTANT: Here's where the the nose of the cat appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the eye of the cat in this image? Please elaborate your answer and explain why. ASSISTANT: This is the selected the eye of the cat.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the the ear of the cat from the image below? ASSISTANT: I've marked the the ear of the cat for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the ear of the cat in this image? Please give some explanation. ASSISTANT: The the ear of the cat portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the cat eye from the provided image. ASSISTANT: Take a look at the cat eye here.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (682, 1024), ['<image>\nPlease highlight the the nose of the cat in this image.', '<image>\nWhat is the eye of the cat in this image? Please elaborate your answer and explain why.', '<image>\nWould you please extract the the ear of the cat from the image below?', '<image>\nWhat is the ear of the cat in this image? Please give some explanation.', '<image>\nSegment the cat eye from the provided image.'], ['the nose of the cat', 'the eye of the cat', 'the ear of the cat', 'the ear of the cat', 'cat eye'], [{'the nose of the cat': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the eye of the cat': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the ear of the cat': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the ear of the cat': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'cat eye': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2009_004643.jpg', tensor([[[0.1939, 0.1939, 0.1768,  ..., 0.2111, 0.1939, 0.1939],
         [0.1939, 0.1939, 0.1768,  ..., 0.2282, 0.2111, 0.2111],
         [0.2111, 0.2111, 0.1939,  ..., 0.2453, 0.2282, 0.2282],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.6254, 0.6254, 0.6078,  ..., 0.6254, 0.6078, 0.6078],
         [0.6254, 0.6254, 0.6078,  ..., 0.6254, 0.6078, 0.6078],
         [0.6254, 0.6254, 0.6254,  ..., 0.6078, 0.5903, 0.5903],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.9668, 0.9842, 1.0365,  ..., 1.1585, 1.1237, 1.1062],
         [0.9668, 0.9842, 1.0365,  ..., 1.1585, 1.1411, 1.1237],
         [0.9668, 0.9842, 1.0191,  ..., 1.1759, 1.1759, 1.1759],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.6092,  1.6238,  1.6092,  ...,  0.7625,  0.7625,  0.7479],
         [ 1.5800,  1.5946,  1.5946,  ...,  0.7917,  0.7771,  0.7625],
         [ 1.5800,  1.5946,  1.5800,  ...,  0.7771,  0.7625,  0.7479],
         ...,
         [-1.1791, -1.0915, -1.3251,  ...,  0.9814, -0.3616, -0.3616],
         [-1.2667, -1.2813, -1.1791,  ..., -0.1572,  0.3975,  0.6749],
         [-1.2959, -1.2521, -1.1791,  ..., -0.1864, -0.7558,  0.0325]],

        [[ 1.4145,  1.4295,  1.4145,  ...,  1.0393,  1.0393,  1.0243],
         [ 1.3845,  1.3995,  1.3995,  ...,  1.0393,  1.0243,  1.0243],
         [ 1.3845,  1.3995,  1.3845,  ...,  1.0393,  1.0243,  1.0093],
         ...,
         [-1.0767, -0.9417, -1.1968,  ...,  0.6792, -0.2813, -0.3414],
         [-1.1668, -1.1668, -1.0617,  ..., -0.5515,  0.2589,  0.7842],
         [-1.1818, -1.1668, -1.0617,  ..., -0.0562, -0.7016,  0.1989]],

        [[ 0.9941,  1.0083,  0.9941,  ...,  1.3780,  1.3638,  1.3496],
         [ 0.9656,  0.9799,  0.9799,  ...,  1.3780,  1.3638,  1.3638],
         [ 0.9656,  0.9799,  0.9656,  ...,  1.3922,  1.3922,  1.3780],
         ...,
         [-0.8688, -0.6981, -0.9399,  ...,  0.7523, -0.0440, -0.2004],
         [-0.9541, -0.9541, -0.7977,  ..., -0.3568,  0.3968,  0.9514],
         [-0.9114, -0.9256, -0.8119,  ...,  0.2120, -0.4422,  0.3826]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the motorbike headlight in this image. ASSISTANT: Below you can see the motorbike headlight.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the motorbike headlight shown in this image. ASSISTANT: You can see the motorbike headlight in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the the wheel of the motorbike in this image. ASSISTANT: This view focuses on the the wheel of the motorbike.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the the head of the person in this image? ASSISTANT: the head of the person.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the person torso from the provided image. ASSISTANT: The person torso is presented here.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (684, 1024), ['<image>\nPlease highlight the motorbike headlight in this image.', '<image>\nDisplay a segmentation mask for the motorbike headlight shown in this image.', '<image>\nPlease highlight the the wheel of the motorbike in this image.', '<image>\nCan you segment the the head of the person in this image?', '<image>\nSegment the person torso from the provided image.'], ['motorbike headlight', 'motorbike headlight', 'the wheel of the motorbike', 'the head of the person', 'person torso'], [{'motorbike headlight': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'motorbike headlight': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the wheel of the motorbike': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the head of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person torso': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000458594.jpg', tensor([[[-1.5185, -1.4672, -1.4158,  ..., -1.9638, -1.9295, -1.9124],
         [-1.6213, -1.5014, -1.3815,  ..., -1.9467, -1.9295, -1.9124],
         [-1.7240, -1.5528, -1.3644,  ..., -1.9124, -1.9295, -1.9295],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.1779, -1.2654, -1.3880,  ..., -1.6506, -1.6331, -1.6331],
         [-1.3004, -1.2654, -1.2304,  ..., -1.6331, -1.6331, -1.6331],
         [-1.4580, -1.2654, -1.0553,  ..., -1.6155, -1.6331, -1.6331],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.2119, -1.2641, -1.2990,  ..., -1.7173, -1.6999, -1.6824],
         [-1.3164, -1.2467, -1.1596,  ..., -1.6824, -1.6824, -1.6650],
         [-1.4036, -1.2293, -0.9853,  ..., -1.6302, -1.6476, -1.6476],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.9018, -1.3981, -1.1499,  ..., -1.3835, -1.4127, -1.4419],
         [-1.0185, -1.3543, -1.1207,  ..., -1.3689, -1.3981, -1.4857],
         [-1.0623, -1.2375, -1.1207,  ..., -1.4419, -1.5003, -1.5003],
         ...,
         [ 1.9011,  1.8865,  1.9157,  ..., -0.7996, -0.8288, -0.8726],
         [ 1.9157,  1.9157,  1.9157,  ..., -0.8726, -1.0477, -1.3689],
         [ 1.9157,  1.9011,  1.9157,  ..., -1.1061, -1.4565, -1.4711]],

        [[-0.4614, -1.2568, -0.9267,  ..., -1.1068, -1.0017, -1.1818],
         [-0.6865, -1.2418, -0.8967,  ..., -1.1068, -1.1668, -1.1518],
         [-0.5515, -1.1218, -1.0167,  ..., -1.0767, -1.2418, -1.2268],
         ...,
         [ 2.0599,  2.0449,  2.0599,  ..., -0.5815, -0.7166, -0.7766],
         [ 2.0599,  2.0599,  2.0599,  ..., -0.6865, -0.9417, -1.3619],
         [ 2.0599,  2.0599,  2.0599,  ..., -1.1068, -1.4219, -1.3319]],

        [[-1.4376, -1.1389, -1.2527,  ..., -1.2954, -1.3665, -1.3238],
         [-1.4233, -1.1958, -1.2954,  ..., -1.1816, -1.2811, -1.3238],
         [-1.4376, -1.2100, -1.2385,  ..., -1.1532, -1.2954, -1.2954],
         ...,
         [ 2.0748,  2.0748,  2.0890,  ..., -0.4990, -0.6412, -0.6128],
         [ 2.1032,  2.1317,  2.1317,  ..., -0.5986, -0.8545, -1.0963],
         [ 2.0890,  2.0890,  2.0890,  ..., -0.9541, -1.1532, -1.1105]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the closest in this image. ASSISTANT: Here's the area for closest.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the zebra behind the full one from the provided image. ASSISTANT: Here's the area for zebra behind the full one.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the the one that is covered but is not covered a lot shown in this image. ASSISTANT: I've marked the the one that is covered but is not covered a lot for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the big zebra close to green grass in this picture? ASSISTANT: The big zebra close to green grass is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is closest zebra in this image? Please elaborate your answer and explain why. ASSISTANT: The region corresponding to the closest zebra is shown.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nPlease segment the closest in this image.', '<image>\nSegment the zebra behind the full one from the provided image.', '<image>\nDisplay a segmentation mask for the the one that is covered but is not covered a lot shown in this image.', '<image>\nCould you isolate the big zebra close to green grass in this picture?', '<image>\nWhat is closest zebra in this image? Please elaborate your answer and explain why.'], ['closest', 'zebra behind the full one', 'the one that is covered but is not covered a lot', 'big zebra close to green grass', 'closest zebra'], [{'closest': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'zebra behind the full one': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the one that is covered but is not covered a lot': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'big zebra close to green grass': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'closest zebra': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000261948.jpg', tensor([[[-1.7240, -1.6555, -1.5870,  ..., -0.3541, -0.3712, -0.3883],
         [-1.7240, -1.6555, -1.6042,  ..., -0.3541, -0.3712, -0.3712],
         [-1.7069, -1.6727, -1.6384,  ..., -0.3369, -0.3541, -0.3541],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.8452, -0.8803, -0.9328,  ..., -0.5651, -0.5651, -0.5651],
         [-0.8627, -0.8803, -0.9328,  ..., -0.5651, -0.5476, -0.5476],
         [-0.8978, -0.8978, -0.9153,  ..., -0.5476, -0.5301, -0.5301],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.7413, -0.7413, -0.7238,  ..., -0.4450, -0.4624, -0.4624],
         [-0.7587, -0.7413, -0.7238,  ..., -0.4450, -0.4450, -0.4450],
         [-0.7761, -0.7587, -0.7238,  ..., -0.4275, -0.4275, -0.4275],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.3835, -1.3397, -1.3251,  ...,  0.0909,  0.0617,  0.1639],
         [-1.3543, -1.3251, -1.2959,  ...,  0.0617, -0.0113,  0.2223],
         [-1.3543, -1.3543, -1.3251,  ...,  0.1201,  0.0033,  0.1639],
         ...,
         [-1.6755, -1.7047, -1.6901,  ...,  0.4997,  0.3537,  0.1639],
         [-1.6609, -1.7047, -1.6609,  ...,  0.3537,  0.4267,  0.3245],
         [-1.6609, -1.7047, -1.6463,  ...,  0.2661,  0.2369,  0.1639]],

        [[-0.4764, -0.4914, -0.5215,  ..., -0.2963, -0.2963, -0.3114],
         [-0.4914, -0.4914, -0.5065,  ..., -0.2513, -0.2963, -0.2363],
         [-0.4764, -0.4764, -0.4914,  ..., -0.2513, -0.2813, -0.2363],
         ...,
         [-1.2118, -1.1368, -1.1668,  ...,  0.6041,  0.4390,  0.2439],
         [-1.2568, -1.2118, -1.1818,  ...,  0.4390,  0.4991,  0.3940],
         [-1.2718, -1.2418, -1.2118,  ...,  0.3490,  0.3040,  0.2139]],

        [[-0.3995, -0.3995, -0.3995,  ..., -0.0582, -0.0582, -0.1009],
         [-0.4137, -0.3995, -0.3711,  ...,  0.0129, -0.0298, -0.0867],
         [-0.3711, -0.3568, -0.3711,  ..., -0.0156, -0.0298, -0.0582],
         ...,
         [-1.0821, -1.0394, -1.0394,  ...,  0.7523,  0.5959,  0.4110],
         [-1.0963, -1.0536, -1.0394,  ...,  0.5817,  0.6528,  0.5532],
         [-1.1247, -1.1105, -1.0963,  ...,  0.4679,  0.4395,  0.3684]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the pipe in this image. ASSISTANT: The region corresponding to the pipe is shown.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the pipe in this image. ASSISTANT: Take a look at the pipe here.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nPlease segment the pipe in this image.', '<image>\nPlease highlight the pipe in this image.'], ['pipe', 'pipe'], [{'pipe': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'pipe': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg')]
[rank1]: Traceback (most recent call last):
[rank1]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 849, in train
[rank1]:     input_dict = next(train_iter)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 124, in __next__
[rank1]:     return next(self.data)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 157, in <genexpr>
[rank1]:     self.data = (x for x in self.dataloader)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 631, in __next__
[rank1]:     data = self._next_data()
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1346, in _next_data
[rank1]:     return self._process_data(data)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1372, in _process_data
[rank1]:     data.reraise()
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/_utils.py", line 705, in reraise
[rank1]:     raise exception
[rank1]: ValueError: Caught ValueError in DataLoader worker process 0.
[rank1]: Original Traceback (most recent call last):
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/worker.py", line 308, in _worker_loop
[rank1]:     data = fetcher.fetch(index)  # type: ignore[possibly-undefined]
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/fetch.py", line 54, in fetch
[rank1]:     return self.collate_fn(data)
[rank1]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/utils/dataset.py", line 53, in collate_fn
[rank1]:     for (
[rank1]: ValueError: too many values to unpack (expected 12)


[rank1]: During handling of the above exception, another exception occurred:

[rank1]: Traceback (most recent call last):
[rank1]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 1282, in <module>
[rank1]:     main(sys.argv[1:])
[rank1]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 743, in main
[rank1]:     train_iter = train(
[rank1]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 852, in train
[rank1]:     input_dict = next(train_iter)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 124, in __next__
[rank1]:     return next(self.data)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 157, in <genexpr>
[rank1]:     self.data = (x for x in self.dataloader)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 631, in __next__
[rank1]:     data = self._next_data()
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1346, in _next_data
[rank1]:     return self._process_data(data)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1372, in _process_data
[rank1]:     data.reraise()
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/_utils.py", line 705, in reraise
[rank1]:     raise exception
[rank1]: ValueError: Caught ValueError in DataLoader worker process 0.
[rank1]: Original Traceback (most recent call last):
[rank1]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 849, in train
[rank1]:     input_dict = next(train_iter)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 124, in __next__
[rank1]:     return next(self.data)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 157, in <genexpr>
[rank1]:     self.data = (x for x in self.dataloader)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 631, in __next__
[rank1]:     data = self._next_data()
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1346, in _next_data
[rank1]:     return self._process_data(data)
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1372, in _process_data
[rank1]:     data.reraise()
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/_utils.py", line 705, in reraise
[rank1]:     raise exception
[rank1]: ValueError: Caught ValueError in DataLoader worker process 0.
[rank1]: Original Traceback (most recent call last):
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/worker.py", line 308, in _worker_loop
[rank1]:     data = fetcher.fetch(index)  # type: ignore[possibly-undefined]
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/fetch.py", line 54, in fetch
[rank1]:     return self.collate_fn(data)
[rank1]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/utils/dataset.py", line 53, in collate_fn
[rank1]:     for (
[rank1]: ValueError: too many values to unpack (expected 12)


[rank1]: During handling of the above exception, another exception occurred:

[rank1]: Traceback (most recent call last):
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/worker.py", line 308, in _worker_loop
[rank1]:     data = fetcher.fetch(index)  # type: ignore[possibly-undefined]
[rank1]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/fetch.py", line 54, in fetch
[rank1]:     return self.collate_fn(data)
[rank1]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/utils/dataset.py", line 53, in collate_fn
[rank1]:     for (
[rank1]: ValueError: too many values to unpack (expected 12)

[2025-10-22 16:11:13,362] [INFO] [utils.py:785:see_memory_usage] Before initializing optimizer states
[2025-10-22 16:11:13,363] [INFO] [utils.py:786:see_memory_usage] MA 27.53 GB         Max_MA 28.01 GB         CA 28.15 GB         Max_CA 28 GB 
[2025-10-22 16:11:13,364] [INFO] [utils.py:793:see_memory_usage] CPU Virtual Memory:  used = 206.8 GB, percent = 20.5%
[2025-10-22 16:11:16,392] [INFO] [utils.py:785:see_memory_usage] After initializing optimizer states
[2025-10-22 16:11:16,393] [INFO] [utils.py:786:see_memory_usage] MA 29.46 GB         Max_MA 30.43 GB         CA 31.06 GB         Max_CA 31 GB 
[2025-10-22 16:11:16,393] [INFO] [utils.py:793:see_memory_usage] CPU Virtual Memory:  used = 208.87 GB, percent = 20.7%
[2025-10-22 16:11:16,393] [INFO] [stage_1_and_2.py:488:__init__] optimizer state initialized
[2025-10-22 16:11:19,426] [INFO] [utils.py:785:see_memory_usage] After initializing ZeRO optimizer
[2025-10-22 16:11:19,427] [INFO] [utils.py:786:see_memory_usage] MA 29.46 GB         Max_MA 29.46 GB         CA 31.06 GB         Max_CA 31 GB 
[2025-10-22 16:11:19,427] [INFO] [utils.py:793:see_memory_usage] CPU Virtual Memory:  used = 211.01 GB, percent = 20.9%
[2025-10-22 16:11:19,433] [INFO] [logging.py:96:log_dist] [Rank 0] DeepSpeed Final Optimizer = adamw
[2025-10-22 16:11:19,433] [INFO] [logging.py:96:log_dist] [Rank 0] DeepSpeed using configured LR scheduler = WarmupDecayLR
[2025-10-22 16:11:19,434] [INFO] [logging.py:96:log_dist] [Rank 0] DeepSpeed LR Scheduler = <deepspeed.runtime.lr_schedules.WarmupDecayLR object at 0x7fb0f9a83640>
[2025-10-22 16:11:19,434] [INFO] [logging.py:96:log_dist] [Rank 0] step=0, skipped=0, lr=[0.0003], mom=[(0.9, 0.95)]
[2025-10-22 16:11:19,437] [INFO] [config.py:960:print] DeepSpeedEngine configuration:
[2025-10-22 16:11:19,437] [INFO] [config.py:964:print]   activation_checkpointing_config  {
    "partition_activations": false, 
    "contiguous_memory_optimization": false, 
    "cpu_checkpointing": false, 
    "number_checkpoints": null, 
    "synchronize_checkpoint_boundary": false, 
    "profile": false
}
[2025-10-22 16:11:19,437] [INFO] [config.py:964:print]   aio_config ................... {'block_size': 1048576, 'queue_depth': 8, 'thread_count': 1, 'single_submit': False, 'overlap_events': True}
[2025-10-22 16:11:19,437] [INFO] [config.py:964:print]   amp_enabled .................. False
[2025-10-22 16:11:19,437] [INFO] [config.py:964:print]   amp_params ................... False
[2025-10-22 16:11:19,437] [INFO] [config.py:964:print]   autotuning_config ............ {
    "enabled": false, 
    "start_step": null, 
    "end_step": null, 
    "metric_path": null, 
    "arg_mappings": null, 
    "metric": "throughput", 
    "model_info": null, 
    "results_dir": "autotuning_results", 
    "exps_dir": "autotuning_exps", 
    "overwrite": true, 
    "fast": true, 
    "start_profile_step": 3, 
    "end_profile_step": 5, 
    "tuner_type": "gridsearch", 
    "tuner_early_stopping": 5, 
    "tuner_num_trials": 50, 
    "model_info_path": null, 
    "mp_size": 1, 
    "max_train_batch_size": null, 
    "min_train_batch_size": 1, 
    "max_train_micro_batch_size_per_gpu": 1.024000e+03, 
    "min_train_micro_batch_size_per_gpu": 1, 
    "num_tuning_micro_batch_sizes": 3
}
[2025-10-22 16:11:19,437] [INFO] [config.py:964:print]   bfloat16_enabled ............. True
[2025-10-22 16:11:19,437] [INFO] [config.py:964:print]   checkpoint_parallel_write_pipeline  False
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   checkpoint_tag_validation_enabled  True
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   checkpoint_tag_validation_fail  False
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   comms_config ................. <deepspeed.comm.config.DeepSpeedCommsConfig object at 0x7faff6d13550>
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   communication_data_type ...... None
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   compression_config ........... {'weight_quantization': {'shared_parameters': {'enabled': False, 'quantizer_kernel': False, 'schedule_offset': 0, 'quantize_groups': 1, 'quantize_verbose': False, 'quantization_type': 'symmetric', 'quantize_weight_in_forward': False, 'rounding': 'nearest', 'fp16_mixed_quantize': False, 'quantize_change_ratio': 0.001}, 'different_groups': {}}, 'activation_quantization': {'shared_parameters': {'enabled': False, 'quantization_type': 'symmetric', 'range_calibration': 'dynamic', 'schedule_offset': 1000}, 'different_groups': {}}, 'sparse_pruning': {'shared_parameters': {'enabled': False, 'method': 'l1', 'schedule_offset': 1000}, 'different_groups': {}}, 'row_pruning': {'shared_parameters': {'enabled': False, 'method': 'l1', 'schedule_offset': 1000}, 'different_groups': {}}, 'head_pruning': {'shared_parameters': {'enabled': False, 'method': 'topk', 'schedule_offset': 1000}, 'different_groups': {}}, 'channel_pruning': {'shared_parameters': {'enabled': False, 'method': 'l1', 'schedule_offset': 1000}, 'different_groups': {}}, 'layer_reduction': {'enabled': False}}
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   curriculum_enabled_legacy .... False
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   curriculum_params_legacy ..... False
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   data_efficiency_config ....... {'enabled': False, 'seed': 1234, 'data_sampling': {'enabled': False, 'num_epochs': 1000, 'num_workers': 0, 'curriculum_learning': {'enabled': False}}, 'data_routing': {'enabled': False, 'random_ltd': {'enabled': False, 'layer_token_lr_schedule': {'enabled': False}}}}
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   data_efficiency_enabled ...... False
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   dataloader_drop_last ......... False
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   disable_allgather ............ False
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   dump_state ................... False
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   dynamic_loss_scale_args ...... None
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   eigenvalue_enabled ........... False
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   eigenvalue_gas_boundary_resolution  1
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   eigenvalue_layer_name ........ bert.encoder.layer
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   eigenvalue_layer_num ......... 0
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   eigenvalue_max_iter .......... 100
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   eigenvalue_stability ......... 1e-06
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   eigenvalue_tol ............... 0.01
[2025-10-22 16:11:19,438] [INFO] [config.py:964:print]   eigenvalue_verbose ........... False
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   elasticity_enabled ........... False
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   flops_profiler_config ........ {
    "enabled": false, 
    "recompute_fwd_factor": 0.0, 
    "profile_step": 1, 
    "module_depth": -1, 
    "top_modules": 1, 
    "detailed": true, 
    "output_file": null
}
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   fp16_auto_cast ............... None
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   fp16_enabled ................. False
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   fp16_master_weights_and_gradients  False
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   global_rank .................. 0
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   grad_accum_dtype ............. None
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   gradient_accumulation_steps .. 10
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   gradient_clipping ............ 1.0
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   gradient_predivide_factor .... 1.0
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   hybrid_engine ................ enabled=False max_out_tokens=512 inference_tp_size=1 release_inference_cache=False pin_parameters=True tp_gather_partition_size=8
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   initial_dynamic_scale ........ 1
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   load_universal_checkpoint .... False
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   loss_scale ................... 1.0
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   memory_breakdown ............. False
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   mics_hierarchial_params_gather  False
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   mics_shard_size .............. -1
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   monitor_config ............... tensorboard=TensorBoardConfig(enabled=False, output_path='', job_name='DeepSpeedJobName') wandb=WandbConfig(enabled=False, group=None, team=None, project='deepspeed') csv_monitor=CSVConfig(enabled=False, output_path='', job_name='DeepSpeedJobName') enabled=False
[2025-10-22 16:11:19,439] [INFO] [config.py:964:print]   nebula_config ................ {
    "enabled": false, 
    "persistent_storage_path": null, 
    "persistent_time_interval": 100, 
    "num_of_version_in_retention": 2, 
    "enable_nebula_load": true, 
    "load_path": null
}
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   optimizer_legacy_fusion ...... False
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   optimizer_name ............... adamw
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   optimizer_params ............. {'lr': 0.0003, 'weight_decay': 0.0, 'betas': (0.9, 0.95)}
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   pipeline ..................... {'stages': 'auto', 'partition': 'best', 'seed_layers': False, 'activation_checkpoint_interval': 0}
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   pld_enabled .................. False
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   pld_params ................... False
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   prescale_gradients ........... False
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   scheduler_name ............... WarmupDecayLR
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   scheduler_params ............. {'total_num_steps': 12500, 'warmup_min_lr': 0, 'warmup_max_lr': 0.0003, 'warmup_num_steps': 100, 'warmup_type': 'linear'}
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   sparse_attention ............. None
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   sparse_gradients_enabled ..... False
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   steps_per_print .............. 10
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   train_batch_size ............. 120
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   train_micro_batch_size_per_gpu  6
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   use_node_local_storage ....... False
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   wall_clock_breakdown ......... False
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   world_size ................... 2
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   zero_allow_untested_optimizer  False
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   zero_config .................. stage=2 contiguous_gradients=True reduce_scatter=True reduce_bucket_size=500000000 allgather_partitions=True allgather_bucket_size=500000000 overlap_comm=True load_from_fp32_weights=True elastic_checkpoint=False offload_param=None offload_optimizer=None sub_group_size=1,000,000,000 cpu_offload_param=None cpu_offload_use_pin_memory=None cpu_offload=None prefetch_bucket_size=50,000,000 param_persistence_threshold=100,000 model_persistence_threshold=sys.maxsize max_live_parameters=1,000,000,000 max_reuse_distance=1,000,000,000 gather_16bit_weights_on_model_save=False stage3_gather_fp16_weights_on_model_save=False ignore_unused_parameters=True legacy_stage1=False round_robin_gradients=False mics_shard_size=-1 mics_hierarchical_params_gather=False memory_efficient_linear=True
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   zero_enabled ................. True
[2025-10-22 16:11:19,440] [INFO] [config.py:964:print]   zero_force_ds_cpu_optimizer .. True
[2025-10-22 16:11:19,441] [INFO] [config.py:964:print]   zero_optimization_stage ...... 2
[2025-10-22 16:11:19,441] [INFO] [config.py:950:print_user_config]   json = {
    "train_micro_batch_size_per_gpu": 6, 
    "gradient_accumulation_steps": 10, 
    "optimizer": {
        "type": "AdamW", 
        "params": {
            "lr": 0.0003, 
            "weight_decay": 0.0, 
            "betas": [0.9, 0.95]
        }
    }, 
    "scheduler": {
        "type": "WarmupDecayLR", 
        "params": {
            "total_num_steps": 1.250000e+04, 
            "warmup_min_lr": 0, 
            "warmup_max_lr": 0.0003, 
            "warmup_num_steps": 100, 
            "warmup_type": "linear"
        }
    }, 
    "fp16": {
        "enabled": false
    }, 
    "bf16": {
        "enabled": true
    }, 
    "gradient_clipping": 1.0, 
    "zero_optimization": {
        "stage": 2, 
        "contiguous_gradients": true, 
        "overlap_comm": true, 
        "reduce_scatter": true, 
        "reduce_bucket_size": 5.000000e+08, 
        "allgather_bucket_size": 5.000000e+08
    }
}
(train) >> AFTER DEEPSPEED
>> (train) Auto-resume from:  ./runs/plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025/plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025_bidirbio_2048_maxlen512_epochs25_bidir_bio_feedback_loop_train_prompt_enc_srates_9_7_7_1_ckpt_model
>> (train) resume exists:  False
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000480489.jpg', tensor([[[0.4508, 0.4337, 0.3994,  ..., 0.4508, 0.5364, 0.6049],
         [0.4851, 0.4679, 0.4508,  ..., 0.4337, 0.5364, 0.6049],
         [0.5193, 0.5193, 0.5193,  ..., 0.4166, 0.5193, 0.5878],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.5728, 0.5553, 0.5203,  ..., 0.5903, 0.6779, 0.7479],
         [0.6078, 0.5903, 0.5728,  ..., 0.5728, 0.6779, 0.7479],
         [0.6429, 0.6429, 0.6429,  ..., 0.5553, 0.6604, 0.7304],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.7228, 0.7054, 0.6705,  ..., 0.8099, 0.8971, 0.9668],
         [0.7576, 0.7576, 0.7402,  ..., 0.7925, 0.8971, 0.9668],
         [0.8099, 0.8274, 0.8274,  ..., 0.7751, 0.8797, 0.9494],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[0.3537, 0.3391, 0.3683,  ..., 0.3829, 0.4559, 0.4413],
         [0.3975, 0.3975, 0.3683,  ..., 0.4267, 0.3975, 0.4267],
         [0.3829, 0.3975, 0.4559,  ..., 0.4559, 0.4267, 0.4851],
         ...,
         [0.6311, 0.5873, 0.6895,  ..., 1.6384, 1.7260, 1.7552],
         [0.6457, 0.7041, 0.8063,  ..., 1.8427, 1.9157, 1.8573],
         [0.2369, 0.5143, 0.6895,  ..., 1.3902, 1.2150, 1.2004]],

        [[0.4390, 0.4240, 0.4540,  ..., 0.4841, 0.5591, 0.5441],
         [0.4841, 0.4841, 0.4390,  ..., 0.5291, 0.4991, 0.5291],
         [0.4691, 0.4841, 0.5441,  ..., 0.5591, 0.5291, 0.5891],
         ...,
         [0.6191, 0.6191, 0.7392,  ..., 1.0694, 1.2044, 1.2194],
         [0.6942, 0.7542, 0.8142,  ..., 1.0994, 1.4596, 1.4596],
         [0.2740, 0.5891, 0.7092,  ..., 1.6547, 1.7297, 1.7297]],

        [[0.7097, 0.6955, 0.7097,  ..., 0.6386, 0.7097, 0.6955],
         [0.7239, 0.7381, 0.6955,  ..., 0.6812, 0.6528, 0.6812],
         [0.6955, 0.7097, 0.7666,  ..., 0.7097, 0.6812, 0.7381],
         ...,
         [0.6812, 0.6812, 0.7666,  ..., 0.6244, 0.5248, 0.5390],
         [0.6812, 0.7381, 0.7666,  ..., 0.3399, 0.6528, 0.6812],
         [0.3115, 0.5675, 0.6670,  ..., 0.5390, 0.5390, 0.4253]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the primary focus of the image? ASSISTANT: The primary focus of the image is the sky full of colorful kites flying near a mountain landscape.</s>USER: Is there a unique or unusually shaped kite in the image? ASSISTANT: Yes, there is a large kite in the shape of the bottom half of a woman in the image.</s>USER: Are there any parachutes in the picture? ASSISTANT: No, the image does not feature parachutes; it is focused on colorful kites flying in the sky.</s>USER: What is the setting for the scene with the kites? ASSISTANT: The setting for the scene with the kites is an outdoor location next to some mountains. This backdrop provides a picturesque and natural environment for flying kites.</s>USER: Are there any decorations on the ground visible in the image? ASSISTANT: Yes, there are assorted decorations lying around on a plastic tarp on the ground. This likely indicates that the area is being used as a gathering spot for kite fliers and spectators to enjoy the event, possibly as part of a kite festival or competition.</s>"], tensor([], size=(0, 427, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (683, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the primary focus of the image? ASSISTANT: The primary focus of the image is the sky full of colorful kites flying near a mountain landscape.</s>USER: Is there a unique or unusually shaped kite in the image? ASSISTANT: Yes, there is a large kite in the shape of the bottom half of a woman in the image.</s>USER: Are there any parachutes in the picture? ASSISTANT: No, the image does not feature parachutes; it is focused on colorful kites flying in the sky.</s>USER: What is the setting for the scene with the kites? ASSISTANT: The setting for the scene with the kites is an outdoor location next to some mountains. This backdrop provides a picturesque and natural environment for flying kites.</s>USER: Are there any decorations on the ground visible in the image? ASSISTANT: Yes, there are assorted decorations lying around on a plastic tarp on the ground. This likely indicates that the area is being used as a gathering spot for kite fliers and spectators to enjoy the event, possibly as part of a kite festival or competition.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the primary focus of the image? ASSISTANT: The primary focus of the image is the sky full of colorful kites flying near a mountain landscape.</s>USER: Is there a unique or unusually shaped kite in the image? ASSISTANT: Yes, there is a large kite in the shape of the bottom half of a woman in the image.</s>USER: Are there any parachutes in the picture? ASSISTANT: No, the image does not feature parachutes; it is focused on colorful kites flying in the sky.</s>USER: What is the setting for the scene with the kites? ASSISTANT: The setting for the scene with the kites is an outdoor location next to some mountains. This backdrop provides a picturesque and natural environment for flying kites.</s>USER: Are there any decorations on the ground visible in the image? ASSISTANT: Yes, there are assorted decorations lying around on a plastic tarp on the ground. This likely indicates that the area is being used as a gathering spot for kite fliers and spectators to enjoy the event, possibly as part of a kite festival or competition.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000420414.jpg', tensor([[[-1.3302, -1.5528, -1.8268,  ..., -2.1008, -1.9124, -1.7754],
         [-1.5528, -1.6727, -1.8268,  ..., -1.6727, -1.1760, -0.8164],
         [-1.8097, -1.8097, -1.8097,  ..., -1.0904, -0.2856,  0.2967],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.2479, -1.4755, -1.7556,  ..., -1.9132, -1.8081, -1.7381],
         [-1.4755, -1.5980, -1.7556,  ..., -1.4580, -1.0728, -0.7752],
         [-1.7381, -1.7381, -1.7381,  ..., -0.8627, -0.1625,  0.3452],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.0898, -1.3164, -1.5953,  ..., -1.6824, -1.7347, -1.7696],
         [-1.3164, -1.4384, -1.5953,  ..., -1.1944, -1.0027, -0.8458],
         [-1.5953, -1.5953, -1.5953,  ..., -0.5670, -0.0964,  0.2173],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.9748, -0.9893, -0.7120,  ..., -0.9310, -0.1864, -0.7996],
         [-0.9893, -0.9164, -0.6244,  ..., -1.1791, -1.4127, -1.5879],
         [-0.9748, -0.7996, -0.7412,  ..., -1.2667, -0.9456, -1.3397],
         ...,
         [ 0.7917,  0.8209,  0.7479,  ...,  0.7917,  0.8063,  0.7917],
         [ 0.8792,  0.8209,  0.8501,  ...,  0.7917,  0.8063,  0.7917],
         [ 0.9084,  0.8938,  0.9084,  ...,  0.7917,  0.7771,  0.7771]],

        [[-0.9867, -0.9867, -0.7616,  ..., -0.8066,  0.0188, -0.6715],
         [-1.0167, -0.9417, -0.7016,  ..., -1.1218, -1.3469, -1.5570],
         [-1.0017, -0.8366, -0.7916,  ..., -1.1968, -0.9567, -1.3019],
         ...,
         [ 1.0243,  1.0393,  0.9793,  ...,  0.8893,  0.9043,  0.8893],
         [ 1.0994,  1.0393,  1.0393,  ...,  0.8893,  0.9043,  0.8893],
         [ 1.0994,  1.0844,  1.0994,  ...,  0.8893,  0.8743,  0.8743]],

        [[-0.9541, -0.9399, -0.8261,  ..., -0.6128,  0.1124, -0.5417],
         [-0.9256, -0.9256, -0.8119,  ..., -0.9967, -1.2100, -1.4376],
         [-0.8830, -0.8688, -0.8545,  ..., -1.0678, -0.9541, -1.0678],
         ...,
         [ 1.0083,  1.0510,  0.9941,  ...,  0.9514,  0.9656,  0.9514],
         [ 1.1078,  1.0510,  1.0652,  ...,  0.9514,  0.9656,  0.9514],
         [ 1.1221,  1.1078,  1.1221,  ...,  0.9514,  0.9372,  0.9372]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the potential benefits of this skiing activity? ASSISTANT: Engaging in skiing activities, as seen in the image where a person is skiing on a snowy slope with snow-covered trees in the background, has multiple potential benefits. First, it serves as an excellent form of exercise that can improve cardiovascular health, strengthen muscles, increase balance and coordination, and contribute to overall physical fitness. Second, skiing can provide stress relief by allowing individuals to immerse themselves in nature, enjoying the fresh air and peaceful surroundings, which can have positive effects on mental health. Third, skiing is a social activity that can be enjoyed with friends or family, creating opportunities for bonding and making lasting memories. Finally, skiing can serve as a means to enjoy the outdoors and experience the beauty of the winter season, which can enhance one's appreciation for the natural environment.</s>"], tensor([], size=(0, 480, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the potential benefits of this skiing activity? ASSISTANT: Engaging in skiing activities, as seen in the image where a person is skiing on a snowy slope with snow-covered trees in the background, has multiple potential benefits. First, it serves as an excellent form of exercise that can improve cardiovascular health, strengthen muscles, increase balance and coordination, and contribute to overall physical fitness. Second, skiing can provide stress relief by allowing individuals to immerse themselves in nature, enjoying the fresh air and peaceful surroundings, which can have positive effects on mental health. Third, skiing is a social activity that can be enjoyed with friends or family, creating opportunities for bonding and making lasting memories. Finally, skiing can serve as a means to enjoy the outdoors and experience the beauty of the winter season, which can enhance one's appreciation for the natural environment.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the potential benefits of this skiing activity? ASSISTANT: Engaging in skiing activities, as seen in the image where a person is skiing on a snowy slope with snow-covered trees in the background, has multiple potential benefits. First, it serves as an excellent form of exercise that can improve cardiovascular health, strengthen muscles, increase balance and coordination, and contribute to overall physical fitness. Second, skiing can provide stress relief by allowing individuals to immerse themselves in nature, enjoying the fresh air and peaceful surroundings, which can have positive effects on mental health. Third, skiing is a social activity that can be enjoyed with friends or family, creating opportunities for bonding and making lasting memories. Finally, skiing can serve as a means to enjoy the outdoors and experience the beauty of the winter season, which can enhance one's appreciation for the natural environment.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2008_006882.jpg', tensor([[[-1.3130, -1.3473, -1.4158,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2959, -1.3302, -1.3815,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2617, -1.2788, -1.3130,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.4397, -0.3712, -0.2342,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.6734,  0.6906,  0.7248,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.2043,  1.2043,  1.1872,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.2850, -0.3200, -0.3901,  ...,  0.0000,  0.0000,  0.0000],
         [-0.2675, -0.3025, -0.3550,  ...,  0.0000,  0.0000,  0.0000],
         [-0.2150, -0.2500, -0.2850,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 0.0126,  0.0651,  0.2227,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.1681,  1.1856,  1.2206,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.7108,  1.7108,  1.6933,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.6302, -1.6476, -1.6824,  ...,  0.0000,  0.0000,  0.0000],
         [-1.6302, -1.6476, -1.6650,  ...,  0.0000,  0.0000,  0.0000],
         [-1.6127, -1.6302, -1.6127,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.8458, -0.7761, -0.6193,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.3045,  0.3393,  0.3742,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.8622,  0.8622,  0.8448,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.6895,  0.9522,  1.0106,  ..., -0.3178, -0.1426, -0.0696],
         [ 0.5435,  0.1931,  0.7041,  ...,  0.1931, -0.0113,  0.1347],
         [ 0.3829, -0.0259, -0.2010,  ..., -0.6536, -0.6974, -0.3178],
         ...,
         [-0.9164, -0.9748, -1.1645,  ..., -0.9310, -0.9893, -1.3105],
         [-0.8580, -1.1937, -1.5879,  ..., -0.7704, -1.0769, -1.0769],
         [-0.9310, -1.7923, -0.9893,  ..., -0.6828, -0.5660, -0.1864]],

        [[ 0.6792,  0.9643,  1.0243,  ...,  0.1839,  0.3040,  0.2139],
         [ 0.5291,  0.1839,  0.7092,  ...,  0.7392,  0.4841,  0.4090],
         [ 0.3790, -0.0412, -0.2213,  ..., -0.2213, -0.2663, -0.0862],
         ...,
         [-0.1163, -0.2063, -0.5365,  ...,  0.3190,  0.1539, -0.2063],
         [-0.0262, -0.3714, -1.0767,  ...,  0.4390, -0.1463, -0.1913],
         [-0.1913, -1.4820, -0.2663,  ...,  0.4540,  0.6491,  1.0694]],

        [[ 0.5675,  0.8377,  0.8803,  ..., -0.3142, -0.2146, -0.2715],
         [ 0.4253,  0.0982,  0.5817,  ...,  0.1551, -0.1293, -0.1862],
         [ 0.2831, -0.1151, -0.2857,  ..., -0.6697, -0.7550, -0.5559],
         ...,
         [-1.1532, -1.2669, -1.3238,  ..., -1.2811, -1.2811, -1.3949],
         [-1.1247, -1.3380, -1.3807,  ..., -1.1816, -1.2527, -1.1532],
         [-1.0110, -1.4802, -1.1247,  ..., -1.1958, -1.1389, -0.6555]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the the torso of the bird in this image? ASSISTANT: The the torso of the bird is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the bird head in this image. ASSISTANT: The bird head portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is bird neck in this image? Please elaborate your answer and explain why. ASSISTANT: Take a look at the bird neck here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the bird foot from the image below? ASSISTANT: Below you can see the bird foot.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the bird wing from the provided image. ASSISTANT: bird wing.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 768), ['<image>\nCan you segment the the torso of the bird in this image?', '<image>\nPlease segment the bird head in this image.', '<image>\nWhat is bird neck in this image? Please elaborate your answer and explain why.', '<image>\nWould you please extract the bird foot from the image below?', '<image>\nSegment the bird wing from the provided image.'], ['the torso of the bird', 'bird head', 'bird neck', 'bird foot', 'bird wing'], [{'the torso of the bird': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'bird head': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'bird neck': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'bird foot': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'bird wing': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/ade20k/images/training/ADE_train_00019821.jpg', tensor([[[-0.9705, -0.9534, -0.9534,  ..., -1.0390, -1.0390, -1.0390],
         [-0.9192, -0.9192, -0.9192,  ..., -1.0390, -1.0390, -1.0390],
         [-0.8678, -0.8849, -0.9020,  ..., -1.0390, -1.0390, -1.0390],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.7227, -0.7052, -0.7052,  ..., -0.9503, -0.9503, -0.9503],
         [-0.6702, -0.6702, -0.6702,  ..., -0.9503, -0.9503, -0.9503],
         [-0.6176, -0.6352, -0.6527,  ..., -0.9503, -0.9503, -0.9503],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.4450, -0.4275, -0.4275,  ..., -0.7936, -0.7936, -0.7936],
         [-0.3927, -0.3927, -0.4101,  ..., -0.7936, -0.7936, -0.7936],
         [-0.3404, -0.3578, -0.3753,  ..., -0.7936, -0.7936, -0.7936],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.7996, -0.7850, -0.8580,  ..., -0.4492, -0.4930, -0.4784],
         [-0.7996, -0.8288, -0.8872,  ..., -0.4930, -0.5076, -0.5076],
         [-0.7558, -0.8288, -0.8580,  ..., -0.5076, -0.5076, -0.5076],
         ...,
         [ 0.3245,  0.3683,  0.3975,  ..., -0.1864, -0.2010, -0.2448],
         [ 0.5435,  0.4997,  0.3683,  ..., -0.2594, -0.2448, -0.2886],
         [ 0.2077,  0.1493,  0.1639,  ..., -0.2886, -0.2740, -0.2740]],

        [[-0.6565, -0.6115, -0.6415,  ..., -0.4464, -0.4914, -0.4914],
         [-0.6415, -0.6565, -0.6715,  ..., -0.4914, -0.5065, -0.5065],
         [-0.5965, -0.6715, -0.6415,  ..., -0.5065, -0.5065, -0.5065],
         ...,
         [ 0.3040,  0.3490,  0.3790,  ..., -0.1613, -0.1763, -0.2063],
         [ 0.5591,  0.5141,  0.3790,  ..., -0.2363, -0.2213, -0.2513],
         [ 0.2439,  0.1989,  0.1989,  ..., -0.2663, -0.2513, -0.2513]],

        [[-0.4422, -0.4279, -0.4848,  ..., -0.3000, -0.3426, -0.3284],
         [-0.4422, -0.4706, -0.4990,  ..., -0.3426, -0.3568, -0.3426],
         [-0.3853, -0.4848, -0.4848,  ..., -0.3568, -0.3568, -0.3426],
         ...,
         [ 0.2262,  0.2688,  0.2973,  ...,  0.0129, -0.0013, -0.0298],
         [ 0.4679,  0.4395,  0.3257,  ..., -0.0582, -0.0440, -0.0724],
         [ 0.1693,  0.1551,  0.1977,  ..., -0.0867, -0.0724, -0.0724]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the rock in this image. ASSISTANT: Take a look at the rock here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the waterfall in this image. ASSISTANT: Displayed here is the waterfall.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ...,  True, False, False],
         [ True,  True,  True,  ...,  True, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[113, 113, 113,  ..., 255, 255, 255],
        [113, 113, 113,  ..., 113, 255, 255],
        [113, 113, 113,  ..., 113, 255, 255],
        ...,
        [255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255]]), (680, 1024), ['<image>\nPlease highlight the rock in this image.', '<image>\nPlease highlight the waterfall in this image.'], ['rock', 'waterfall'], [{'rock': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'waterfall': tensor([[ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ...,  True, False, False],
        [ True,  True,  True,  ...,  True, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000422061.jpg', tensor([[[2.2489, 2.2489, 2.2489,  ..., 2.0263, 1.9578, 1.9064],
         [2.2489, 2.2489, 2.2489,  ..., 2.0092, 1.9920, 1.9749],
         [2.2489, 2.2489, 2.2489,  ..., 2.0092, 2.0263, 2.0777],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.4286, 2.4286, 2.4286,  ..., 2.3585, 2.3060, 2.2535],
         [2.4286, 2.4286, 2.4286,  ..., 2.3060, 2.3060, 2.3060],
         [2.4286, 2.4286, 2.4286,  ..., 2.2535, 2.3060, 2.3585],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.6400, 2.6400, 2.6400,  ..., 2.6400, 2.5703, 2.5006],
         [2.6400, 2.6400, 2.6400,  ..., 2.6226, 2.5529, 2.5006],
         [2.6400, 2.6400, 2.6400,  ..., 2.5877, 2.5354, 2.5006],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.9303,  1.9303,  1.9303,  ...,  1.9303,  1.9303,  1.9303],
         [ 1.9303,  1.9303,  1.9303,  ...,  1.9303,  1.9303,  1.9303],
         [ 1.9303,  1.9303,  1.9303,  ...,  1.9303,  1.9303,  1.9303],
         ...,
         [-0.2156, -0.1426,  0.3683,  ...,  0.2807,  0.1347,  0.0909],
         [ 0.0909,  0.6457,  0.5581,  ...,  0.5289,  0.1785,  0.0325],
         [ 0.5727,  0.7187,  0.4559,  ...,  0.1493,  0.4413, -0.0113]],

        [[ 2.0749,  2.0749,  2.0749,  ...,  2.0749,  2.0749,  2.0749],
         [ 2.0749,  2.0749,  2.0749,  ...,  2.0749,  2.0749,  2.0749],
         [ 2.0749,  2.0749,  2.0749,  ...,  2.0749,  2.0749,  2.0749],
         ...,
         [-0.5215, -0.2663,  0.1239,  ..., -0.1613, -0.2663, -0.2663],
         [-0.0412,  0.4390,  0.2589,  ...,  0.0638, -0.2363, -0.2963],
         [ 0.3940,  0.3040,  0.0789,  ..., -0.2213,  0.0638, -0.4164]],

        [[ 2.1459,  2.1459,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         [ 2.1459,  2.1459,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         [ 2.1459,  2.1459,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         ...,
         [-0.2715, -0.1578,  0.0698,  ..., -0.0298, -0.1009, -0.0867],
         [-0.0298,  0.5248,  0.4395,  ...,  0.1693, -0.0867, -0.1435],
         [ 0.5675,  0.3684,  0.1124,  ..., -0.1435,  0.1409, -0.2573]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is a white horse carrying a man in this image? Please elaborate your answer and explain why. ASSISTANT: This view focuses on the a white horse carrying a man.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the a white horse near by dark brown horse shown in this image. ASSISTANT: You can see the a white horse near by dark brown horse in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the woman on horse in this picture. ASSISTANT: You can see the woman on horse in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the woman riding dark brown horse shown in this image. ASSISTANT: The region corresponding to the woman riding dark brown horse is shown.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (666, 1024), ['<image>\nWhat is a white horse carrying a man in this image? Please elaborate your answer and explain why.', '<image>\nDisplay a segmentation mask for the a white horse near by dark brown horse shown in this image.', '<image>\nPlease point out the woman on horse in this picture.', '<image>\nDisplay a segmentation mask for the woman riding dark brown horse shown in this image.'], ['a white horse carrying a man', 'a white horse near by dark brown horse', 'woman on horse', 'woman riding dark brown horse'], [{'a white horse carrying a man': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'a white horse near by dark brown horse': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'woman on horse': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'woman riding dark brown horse': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2008_001746.jpg', tensor([[[ 0.0741,  0.0912,  0.1083,  ..., -0.7650, -0.5082, -0.3883],
         [ 0.1939,  0.2111,  0.2111,  ..., -0.7137, -0.4911, -0.3883],
         [ 0.4337,  0.4337,  0.4337,  ..., -0.6281, -0.4397, -0.3712],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.1877,  0.2052,  0.2227,  ..., -0.7752, -0.6001, -0.5126],
         [ 0.3102,  0.3277,  0.3277,  ..., -0.7227, -0.5651, -0.4951],
         [ 0.5553,  0.5553,  0.5553,  ..., -0.6176, -0.4951, -0.4426],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.5485,  0.5485,  0.5659,  ..., -0.4973, -0.3927, -0.3404],
         [ 0.6705,  0.6705,  0.6705,  ..., -0.4798, -0.3927, -0.3404],
         [ 0.9145,  0.8971,  0.8971,  ..., -0.4450, -0.3753, -0.3578],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.3543, -1.2083, -1.1353,  ...,  0.5581,  0.6895,  0.6457],
         [-1.3251, -1.2229, -0.9456,  ...,  0.4267,  0.4413,  0.2661],
         [-1.3397, -1.2375, -1.0185,  ..., -0.2156, -0.2594, -0.2594],
         ...,
         [-0.0988, -0.0113,  0.0763,  ..., -0.4200, -0.4054, -0.3470],
         [ 1.2734,  1.3026,  1.3464,  ..., -0.4200, -0.4492, -0.4346],
         [ 1.1858,  1.2296,  1.2004,  ..., -0.4638, -0.4200, -0.4346]],

        [[-1.2118, -1.1068, -1.0617,  ...,  0.4991,  0.6041,  0.5591],
         [-1.1968, -1.1218, -0.9117,  ...,  0.3790,  0.3940,  0.2139],
         [-1.1968, -1.1518, -1.0317,  ..., -0.2513, -0.2813, -0.2813],
         ...,
         [ 0.1839,  0.2740,  0.3490,  ..., -0.1913, -0.1613, -0.1012],
         [ 1.5946,  1.6247,  1.6547,  ..., -0.1913, -0.2063, -0.1763],
         [ 1.4446,  1.4896,  1.4746,  ..., -0.2213, -0.2063, -0.1913]],

        [[-0.7834, -0.7266, -0.7408,  ...,  0.4821,  0.6101,  0.6244],
         [-0.7834, -0.7692, -0.6412,  ...,  0.4110,  0.4253,  0.2973],
         [-0.7977, -0.7977, -0.7550,  ..., -0.1435, -0.1862, -0.1578],
         ...,
         [ 0.5532,  0.5959,  0.6670,  ...,  0.3826,  0.3826,  0.3684],
         [ 1.8046,  1.8331,  1.8757,  ...,  0.3257,  0.3257,  0.2973],
         [ 1.7620,  1.8046,  1.7335,  ...,  0.2688,  0.2973,  0.2688]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the the head of the person in this image? ASSISTANT: The the head of the person is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the person arm in this image? ASSISTANT: Here is the person arm you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the person hair in this image. ASSISTANT: Sure — this shows the person hair.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the person ear in this image? ASSISTANT: This view focuses on the person ear.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the the wheel of the car in this image? ASSISTANT: You can see the the wheel of the car in this frame.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nCan you segment the the head of the person in this image?', '<image>\nCan you segment the person arm in this image?', '<image>\nPlease segment the person hair in this image.', '<image>\nCan you segment the person ear in this image?', '<image>\nCan you segment the the wheel of the car in this image?'], ['the head of the person', 'person arm', 'person hair', 'person ear', 'the wheel of the car'], [{'the head of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person arm': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person hair': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person ear': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the wheel of the car': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg')]
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000355210.jpg', tensor([[[-1.2788, -1.3130, -1.3473,  ..., -1.4158, -1.3815, -1.3644],
         [-1.2788, -1.3130, -1.3473,  ..., -1.4158, -1.3815, -1.3644],
         [-1.2959, -1.3130, -1.3302,  ..., -1.4158, -1.3987, -1.3815],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.1954, -1.2304, -1.2654,  ..., -1.3354, -1.3004, -1.2829],
         [-1.1954, -1.2304, -1.2654,  ..., -1.3354, -1.3004, -1.2829],
         [-1.2129, -1.2304, -1.2479,  ..., -1.3354, -1.3179, -1.3004],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.0550, -1.0898, -1.1247,  ..., -1.1770, -1.1421, -1.1247],
         [-1.0550, -1.0898, -1.1247,  ..., -1.1770, -1.1421, -1.1247],
         [-1.0724, -1.0898, -1.1073,  ..., -1.1770, -1.1596, -1.1421],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.9748, -0.9748, -0.9602,  ..., -1.0039, -1.0331, -1.0477],
         [-0.9456, -0.9893, -0.9602,  ..., -1.0039, -1.0477, -1.0185],
         [-0.9310, -0.9456, -0.9456,  ..., -1.0185, -1.0039, -1.0185],
         ...,
         [-0.9164, -0.8872, -0.8726,  ..., -1.7047, -1.7047, -1.6755],
         [-0.9164, -0.9018, -0.8872,  ..., -1.6755, -1.6901, -1.6901],
         [-0.9310, -0.9018, -0.9310,  ..., -1.6755, -1.6901, -1.7047]],

        [[-0.9267, -0.9117, -0.8967,  ..., -0.9417, -0.9717, -1.0017],
         [-0.9117, -0.9117, -0.8967,  ..., -0.9417, -0.9867, -0.9867],
         [-0.9117, -0.9117, -0.9117,  ..., -0.9567, -0.9417, -0.9717],
         ...,
         [-0.8366, -0.8366, -0.8216,  ..., -1.6621, -1.6621, -1.6470],
         [-0.8516, -0.8366, -0.8216,  ..., -1.6320, -1.6470, -1.6621],
         [-0.8666, -0.8516, -0.8666,  ..., -1.6320, -1.6470, -1.6771]],

        [[-0.7834, -0.8119, -0.8403,  ..., -0.8545, -0.8545, -0.8261],
         [-0.7550, -0.8261, -0.8403,  ..., -0.8688, -0.8688, -0.8119],
         [-0.7692, -0.8261, -0.8688,  ..., -0.8688, -0.8261, -0.8119],
         ...,
         [-0.5986, -0.5559, -0.5417,  ..., -1.3949, -1.3949, -1.3522],
         [-0.6128, -0.5701, -0.5417,  ..., -1.3665, -1.3522, -1.3380],
         [-0.6128, -0.5701, -0.5844,  ..., -1.3665, -1.3522, -1.3380]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the nearest chair in this image? ASSISTANT: The nearest chair portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is brown chair in this image? Please elaborate your answer and explain why. ASSISTANT: The brown chair portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the small furniture covering two chairs in this picture. ASSISTANT: Displayed here is the small furniture covering two chairs.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the bar chair next to doorway from the provided image. ASSISTANT: Here's the area for bar chair next to doorway.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the chair that is darker not a barstool from the image below? ASSISTANT: I've marked the chair that is darker not a barstool for you.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nCan you segment the nearest chair in this image?', '<image>\nWhat is brown chair in this image? Please elaborate your answer and explain why.', '<image>\nPlease point out the small furniture covering two chairs in this picture.', '<image>\nSegment the bar chair next to doorway from the provided image.', '<image>\nWould you please extract the chair that is darker not a barstool from the image below?'], ['nearest chair', 'brown chair', 'small furniture covering two chairs', 'bar chair next to doorway', 'chair that is darker not a barstool'], [{'nearest chair': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'brown chair': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'small furniture covering two chairs': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'bar chair next to doorway': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'chair that is darker not a barstool': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/reason_seg/ReasonSeg/train/4412711097_09ac9e1445_o.jpg', tensor([[[ 2.2489,  2.2489,  2.2489,  ...,  0.5707,  0.5878,  0.5878],
         [ 2.2489,  2.2489,  2.2489,  ...,  0.5707,  0.5878,  0.5878],
         [ 2.2489,  2.2489,  2.2489,  ...,  0.5707,  0.5878,  0.5878],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.5182,  1.5182,  1.5182,  ..., -0.0399, -0.0224, -0.0224],
         [ 1.5182,  1.5182,  1.5182,  ..., -0.0399, -0.0224, -0.0224],
         [ 1.4832,  1.4832,  1.5007,  ..., -0.0399, -0.0224, -0.0224],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.7337,  1.7337,  1.7337,  ...,  0.3568,  0.3742,  0.3742],
         [ 1.7337,  1.7337,  1.7337,  ...,  0.3568,  0.3742,  0.3742],
         [ 1.7163,  1.7163,  1.7163,  ...,  0.3568,  0.3742,  0.3742],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.9303,  1.9303,  1.9303,  ...,  1.2880,  1.2588,  1.2588],
         [ 1.9303,  1.9303,  1.9303,  ...,  1.1712,  1.1274,  1.1420],
         [ 1.9303,  1.9303,  1.9303,  ...,  1.1128,  1.0544,  1.0690],
         ...,
         [ 0.3683,  0.2807,  0.3099,  ...,  1.9303,  1.9303,  1.9303],
         [ 0.5727,  0.4267,  0.4559,  ...,  1.9303,  1.9303,  1.9303],
         [ 0.7041,  0.5289,  0.4997,  ...,  1.9303,  1.9303,  1.9303]],

        [[ 1.2344,  1.2795,  1.3395,  ...,  0.4691,  0.4991,  0.5591],
         [ 1.2645,  1.3245,  1.3695,  ...,  0.3490,  0.3790,  0.4390],
         [ 1.2795,  1.3395,  1.3695,  ...,  0.3040,  0.3040,  0.3490],
         ...,
         [-0.1913, -0.2663, -0.1913,  ...,  1.2344,  1.2795,  1.2945],
         [-0.1913, -0.2663, -0.1763,  ...,  1.1444,  1.1894,  1.2044],
         [-0.1463, -0.2663, -0.1763,  ...,  1.0994,  1.0994,  1.1144]],

        [[ 1.3496,  1.3922,  1.4491,  ...,  0.4110,  0.3968,  0.4537],
         [ 1.2074,  1.2643,  1.3211,  ...,  0.3115,  0.3542,  0.4110],
         [ 1.1505,  1.2074,  1.2500,  ...,  0.3542,  0.3399,  0.4253],
         ...,
         [-0.0724, -0.1435, -0.0582,  ...,  1.1932,  1.2216,  1.2074],
         [-0.0582, -0.1293, -0.0440,  ...,  1.1789,  1.2216,  1.2216],
         [-0.0582, -0.1293, -0.0298,  ...,  1.1932,  1.2216,  1.2216]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDogs often enjoy playing with toys, especially those that they can interact with using their paws. What in the picture could the dog use its paws to interact with and play?. Please elaborate your answer and explain why. ASSISTANT: Here is something designed for dogs to interact with using their paws and enjoy as a toy.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat items in the picture might a dog utilize its paws to play with and interact with, given that dogs often enjoy playing with toys in this manner?. Please give some explanation. ASSISTANT: Dogs could use their paws to interact with and play with this.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDogs often enjoy playing with toys, particularly those that they can engage with utilizing their paws. What item in the picture might the dog use its paws to interact with and play?. Please respond with segmentation mask. ASSISTANT: Dogs could use their paws to interact with and play with this.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat in the picture could be utilized by a dog to play and interact using its paws?. Please elaborate your answer and explain why. ASSISTANT: Dogs could use their paws to interact with and play with this, especially if it's a toy they enjoy.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat might a dog use its paws to interact with and play, based on the objects in the picture?. Please respond with segmentation mask. ASSISTANT: Dogs might use their paws to interact with and play with this.</s>"], tensor([[[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]],

        [[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]],

        [[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]],

        [[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]],

        [[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]]]), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nDogs often enjoy playing with toys, especially those that they can interact with using their paws. What in the picture could the dog use its paws to interact with and play?. Please elaborate your answer and explain why.', '<image>\nWhat items in the picture might a dog utilize its paws to play with and interact with, given that dogs often enjoy playing with toys in this manner?. Please give some explanation.', '<image>\nDogs often enjoy playing with toys, particularly those that they can engage with utilizing their paws. What item in the picture might the dog use its paws to interact with and play?. Please respond with segmentation mask.', '<image>\nWhat in the picture could be utilized by a dog to play and interact using its paws?. Please elaborate your answer and explain why.', '<image>\nWhat might a dog use its paws to interact with and play, based on the objects in the picture?. Please respond with segmentation mask.'], ['Here is something designed for dogs to interact with using their paws and enjoy as a toy.', 'Dogs could use their paws to interact with and play with this.', 'Dogs could use their paws to interact with and play with this.', "Dogs could use their paws to interact with and play with this, especially if it's a toy they enjoy.", 'Dogs might use their paws to interact with and play with this.'], [{'Dogs often enjoy playing with toys, especially those that they can interact with using their paws. What in the picture could the dog use its paws to interact with and play?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}, {'What items in the picture might a dog utilize its paws to play with and interact with, given that dogs often enjoy playing with toys in this manner?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}, {'Dogs often enjoy playing with toys, particularly those that they can engage with utilizing their paws. What item in the picture might the dog use its paws to interact with and play?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}, {'What in the picture could be utilized by a dog to play and interact using its paws?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}, {'What might a dog use its paws to interact with and play, based on the objects in the picture?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}], False, 'reason_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000282691.jpg', tensor([[[-1.2103, -1.2103, -1.2103,  ..., -0.1143, -0.1143, -0.1143],
         [-1.2103, -1.2103, -1.2103,  ..., -0.0972, -0.1143, -0.1143],
         [-1.2274, -1.2274, -1.2274,  ..., -0.0801, -0.0972, -0.0972],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.0378, -1.0728, -1.1078,  ...,  0.2052,  0.1877,  0.1702],
         [-1.0553, -1.0728, -1.1078,  ...,  0.2227,  0.1877,  0.1702],
         [-1.0728, -1.0903, -1.1078,  ...,  0.2402,  0.2052,  0.1877],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.0550, -1.0550, -1.0550,  ...,  0.4614,  0.4788,  0.4788],
         [-1.0550, -1.0550, -1.0376,  ...,  0.4962,  0.4962,  0.4962],
         [-1.0376, -1.0376, -1.0201,  ...,  0.5311,  0.5311,  0.5311],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.4054, -0.3908, -0.3616,  ...,  0.6019,  0.6165,  0.6165],
         [-0.3616, -0.3762, -0.3470,  ...,  0.7187,  0.7187,  0.7041],
         [-0.3762, -0.3762, -0.3616,  ...,  0.7625,  0.7917,  0.7625],
         ...,
         [-1.3689, -1.3835, -1.3397,  ..., -1.0331, -0.9893, -0.9310],
         [-1.3397, -1.3397, -1.3689,  ..., -0.9310, -0.9456, -1.0331],
         [-1.3689, -1.4565, -1.4857,  ..., -0.9310, -0.9893, -1.1207]],

        [[-0.2963, -0.2663, -0.2513,  ...,  0.8442,  0.8743,  0.8442],
         [-0.2663, -0.2513, -0.2513,  ...,  0.9043,  0.9193,  0.9193],
         [-0.2663, -0.2663, -0.2663,  ...,  0.9493,  0.9643,  0.9643],
         ...,
         [-1.2718, -1.2268, -1.2568,  ..., -0.8066, -0.7466, -0.7016],
         [-1.2568, -1.2568, -1.2118,  ..., -0.7016, -0.7166, -0.8816],
         [-1.2418, -1.2869, -1.2268,  ..., -0.7316, -0.8516, -1.0317]],

        [[-0.3568, -0.3568, -0.3568,  ...,  0.9799,  1.0083,  1.0083],
         [-0.3711, -0.3568, -0.3426,  ...,  1.0794,  1.0936,  1.0936],
         [-0.3853, -0.3711, -0.3284,  ...,  1.1221,  1.1505,  1.1363],
         ...,
         [-1.2243, -1.1532, -1.1105,  ..., -0.5844, -0.4848, -0.4137],
         [-1.1389, -1.1247, -1.0678,  ..., -0.4279, -0.3568, -0.4137],
         [-1.0963, -1.0678, -0.9541,  ..., -0.3568, -0.4279, -0.5986]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat considerations should a guest make when choosing to book this hotel room? ASSISTANT: When choosing to book this hotel room, a guest should consider the size and layout of the space, the sleeping accommodations, and available amenities. In the image, the room features two beds, making it suitable for guests who require separate sleeping arrangements or are traveling with friends or family. Additionally, there is a flat-screen TV, dining table, and a chair as part of the room's furnishings, which provide basic entertainment and functional needs. The layout has the entrance door leading directly into the sleeping area, which might be an important factor to consider for guests concerned about privacy or noise coming from the hallway. The guest should also think about the cleanliness of the room, which appears to be well-maintained in the image. It is important to evaluate how suitable the room is not only for their personal preferences but also for their intended duration of stay, budget, and any specific requirements they may have for a comfortable and convenient hotel experience.</s>"], tensor([], size=(0, 480, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat considerations should a guest make when choosing to book this hotel room? ASSISTANT: When choosing to book this hotel room, a guest should consider the size and layout of the space, the sleeping accommodations, and available amenities. In the image, the room features two beds, making it suitable for guests who require separate sleeping arrangements or are traveling with friends or family. Additionally, there is a flat-screen TV, dining table, and a chair as part of the room's furnishings, which provide basic entertainment and functional needs. The layout has the entrance door leading directly into the sleeping area, which might be an important factor to consider for guests concerned about privacy or noise coming from the hallway. The guest should also think about the cleanliness of the room, which appears to be well-maintained in the image. It is important to evaluate how suitable the room is not only for their personal preferences but also for their intended duration of stay, budget, and any specific requirements they may have for a comfortable and convenient hotel experience.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat considerations should a guest make when choosing to book this hotel room? ASSISTANT: When choosing to book this hotel room, a guest should consider the size and layout of the space, the sleeping accommodations, and available amenities. In the image, the room features two beds, making it suitable for guests who require separate sleeping arrangements or are traveling with friends or family. Additionally, there is a flat-screen TV, dining table, and a chair as part of the room's furnishings, which provide basic entertainment and functional needs. The layout has the entrance door leading directly into the sleeping area, which might be an important factor to consider for guests concerned about privacy or noise coming from the hallway. The guest should also think about the cleanliness of the room, which appears to be well-maintained in the image. It is important to evaluate how suitable the room is not only for their personal preferences but also for their intended duration of stay, budget, and any specific requirements they may have for a comfortable and convenient hotel experience.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/ade20k/images/training/ADE_train_00020158.jpg', tensor([[[ 1.1872,  1.1872,  1.1872,  ..., -0.8678, -0.8678, -0.8678],
         [ 1.1872,  1.1872,  1.1872,  ..., -0.8678, -0.8678, -0.8678],
         [ 1.1872,  1.1872,  1.1872,  ..., -0.8507, -0.8507, -0.8507],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.0630,  1.0630,  1.0630,  ..., -1.0553, -1.0553, -1.0553],
         [ 1.0630,  1.0630,  1.0630,  ..., -1.0553, -1.0553, -1.0553],
         [ 1.0630,  1.0630,  1.0630,  ..., -1.0378, -1.0378, -1.0378],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.5071,  1.5071,  1.5071,  ..., -0.9678, -0.9678, -0.9678],
         [ 1.5071,  1.5071,  1.5071,  ..., -0.9678, -0.9678, -0.9678],
         [ 1.5071,  1.5071,  1.5071,  ..., -0.9504, -0.9504, -0.9504],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.1493,  0.2807,  0.9230,  ..., -0.4930, -0.5368, -0.4784],
         [ 0.4121,  0.3245,  0.2953,  ..., -0.4930, -0.5514, -0.4784],
         [ 0.7479,  0.3537,  0.1931,  ..., -0.4930, -0.5514, -0.4784],
         ...,
         [ 1.6822,  1.7260,  1.7844,  ..., -0.6098, -0.5806, -0.4784],
         [ 1.7990,  1.7698,  1.7698,  ..., -0.4930, -0.4492, -0.4054],
         [ 1.8427,  1.7552,  1.7552,  ..., -0.2886, -0.2302, -0.2156]],

        [[ 0.3490,  0.4841,  1.1594,  ..., -0.7166, -0.7616, -0.7016],
         [ 0.6191,  0.5291,  0.4841,  ..., -0.7166, -0.7766, -0.7016],
         [ 0.9193,  0.4841,  0.3340,  ..., -0.7166, -0.7766, -0.7016],
         ...,
         [ 1.6247,  1.6997,  1.7597,  ..., -0.6115, -0.7016, -0.7016],
         [ 1.7597,  1.7447,  1.7597,  ..., -0.5815, -0.6265, -0.7016],
         [ 1.8047,  1.7447,  1.7447,  ..., -0.4464, -0.4614, -0.5815]],

        [[-0.1293, -0.0867,  0.4537,  ..., -0.6555, -0.6981, -0.6412],
         [ 0.1693, -0.0298, -0.1293,  ..., -0.6555, -0.7123, -0.6412],
         [ 0.5390,  0.0271, -0.2004,  ..., -0.6555, -0.7123, -0.6412],
         ...,
         [ 1.7051,  1.7620,  1.8188,  ..., -0.4564, -0.5844, -0.8545],
         [ 1.8615,  1.8046,  1.8188,  ..., -0.5417, -0.6981, -0.9967],
         [ 1.9042,  1.8188,  1.8046,  ..., -0.4990, -0.6555, -0.9825]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the wall in this image. ASSISTANT: Below you can see the wall.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the bed in this image. ASSISTANT: This is the selected bed.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the painting from the image below? ASSISTANT: Below you can see the painting.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the pillow from the image below? ASSISTANT: Sure — this shows the pillow.</s>"], tensor([[[False, False, False,  ...,  True,  True, False],
         [ True,  True,  True,  ...,  True,  True, False],
         [ True,  True,  True,  ...,  True,  True, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[255, 255, 255,  ...,   0,   0, 255],
        [  0,   0,   0,  ...,   0,   0, 255],
        [  0,   0,   0,  ...,   0,   0, 255],
        ...,
        [  7,   7,   7,  ..., 255, 255, 255],
        [  7,   7,   7,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255]]), (768, 1024), ['<image>\nPlease highlight the wall in this image.', '<image>\nPlease highlight the bed in this image.', '<image>\nWould you please extract the painting from the image below?', '<image>\nWould you please extract the pillow from the image below?'], ['wall', 'bed', 'painting', 'pillow'], [{'wall': tensor([[False, False, False,  ...,  True,  True, False],
        [ True,  True,  True,  ...,  True,  True, False],
        [ True,  True,  True,  ...,  True,  True, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'bed': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'painting': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'pillow': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/saiapr_tc-12/31/images/31767.jpg', tensor([[[-0.9020, -0.9020, -0.8849,  ...,  0.4508, -0.2684, -0.5767],
         [-0.9705, -0.9534, -0.9192,  ...,  0.4508, -0.2856, -0.6109],
         [-1.1247, -1.0904, -1.0048,  ...,  0.4679, -0.3369, -0.6965],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.8277, -0.8102, -0.7927,  ...,  0.3452, -0.3901, -0.7052],
         [-0.8978, -0.8627, -0.8277,  ...,  0.3627, -0.4076, -0.7402],
         [-1.0553, -1.0028, -0.9153,  ...,  0.3803, -0.4426, -0.7927],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.5147, -0.4973, -0.4798,  ...,  0.7925,  0.0605, -0.2532],
         [-0.5844, -0.5495, -0.5147,  ...,  0.8099,  0.0431, -0.2707],
         [-0.7413, -0.6890, -0.6018,  ...,  0.8448,  0.0256, -0.3404],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.7704, -0.5076, -0.7558,  ..., -0.5368, -0.5514, -0.5952],
         [-0.8142, -0.8726, -0.9893,  ..., -0.6098, -0.6536, -0.7120],
         [-0.6098, -0.7704, -0.7266,  ..., -0.6828, -0.6974, -0.7558],
         ...,
         [ 0.3245,  0.3245,  0.3391,  ...,  0.2223,  0.2223,  0.2661],
         [ 0.3975,  0.3829,  0.3683,  ...,  0.2369,  0.2369,  0.2661],
         [ 0.4121,  0.4121,  0.3683,  ...,  0.2369,  0.2223,  0.2223]],

        [[-0.7316, -0.3264, -0.6865,  ..., -0.4014, -0.4014, -0.4314],
         [-0.8216, -0.7316, -0.8366,  ..., -0.4614, -0.4914, -0.5365],
         [-0.7316, -0.8216, -0.7166,  ..., -0.5065, -0.5365, -0.5965],
         ...,
         [ 0.8893,  0.9043,  0.9193,  ...,  0.7842,  0.7692,  0.7992],
         [ 0.9343,  0.9493,  0.9343,  ...,  0.7992,  0.7542,  0.7842],
         [ 0.9193,  0.9193,  0.9043,  ...,  0.7842,  0.7842,  0.7542]],

        [[-0.5275, -0.1578, -0.4848,  ..., -0.2004, -0.2146, -0.2715],
         [-0.6412, -0.4422, -0.5417,  ..., -0.3000, -0.3142, -0.3711],
         [-0.5559, -0.5559, -0.4564,  ..., -0.3426, -0.3426, -0.3995],
         ...,
         [ 0.3968,  0.4253,  0.4395,  ...,  0.3115,  0.2973,  0.3542],
         [ 0.4679,  0.4679,  0.4679,  ...,  0.3399,  0.2973,  0.3399],
         [ 0.4679,  0.4395,  0.4110,  ...,  0.3684,  0.3399,  0.3257]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is green floor of court not the net in this image? Please elaborate your answer and explain why. ASSISTANT: green floor of court not the net.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the court , not net in this image? ASSISTANT: Below you can see the court , not net.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nWhat is green floor of court not the net in this image? Please elaborate your answer and explain why.', '<image>\nCan you segment the court , not net in this image?'], ['green floor of court not the net', 'court , not net'], [{'green floor of court not the net': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1]], dtype=torch.uint8)}, {'court , not net': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000442886.jpg', tensor([[[-1.7583, -1.8439, -1.9467,  ..., -1.5699, -1.7754, -1.9124],
         [-1.8097, -1.8953, -1.9809,  ..., -1.7412, -1.8782, -1.9467],
         [-1.8953, -1.9467, -2.0152,  ..., -1.9638, -1.9980, -1.9980],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.7206, -1.7731, -1.8431,  ..., -1.5980, -1.7731, -1.8957],
         [-1.8081, -1.8606, -1.9132,  ..., -1.7556, -1.8606, -1.9307],
         [-1.9307, -1.9657, -1.9832,  ..., -1.9482, -1.9657, -1.9832],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.7870, -1.7870, -1.8044,  ..., -1.5779, -1.6999, -1.8044],
         [-1.6824, -1.7347, -1.7870,  ..., -1.6650, -1.7522, -1.8044],
         [-1.5604, -1.6476, -1.7522,  ..., -1.7696, -1.7870, -1.7870],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.9310, -0.8726, -0.8726,  ..., -1.5003, -1.5879, -1.6609],
         [-1.0623, -0.8288, -0.7120,  ..., -1.5733, -1.6463, -1.6609],
         [-1.5149, -1.2959, -0.9456,  ..., -1.6025, -1.6755, -1.6463],
         ...,
         [-0.0696,  0.3099,  0.5289,  ...,  0.7333, -0.1718, -1.6755],
         [-0.0988,  0.2661,  0.5873,  ...,  0.6895,  0.1201, -1.6171],
         [ 0.2077,  0.4997,  0.8938,  ...,  0.6019,  0.4121, -1.4857]],

        [[-0.9867, -1.0017, -0.9117,  ..., -1.6621, -1.6921, -1.7071],
         [-1.0467, -0.7166, -0.5665,  ..., -1.6921, -1.7071, -1.7071],
         [-1.5120, -1.4069, -0.9417,  ..., -1.6921, -1.6921, -1.7071],
         ...,
         [-1.4519, -0.8666, -0.4314,  ..., -0.3714, -0.5365, -1.6621],
         [-1.4669, -0.9267, -0.2663,  ..., -0.3114, -0.3264, -1.6320],
         [-1.1068, -0.5665,  0.2740,  ..., -0.4164, -0.1163, -1.5270]],

        [[-1.4233, -1.4233, -1.4233,  ..., -1.4660, -1.4518, -1.4802],
         [-1.4233, -1.4233, -1.4518,  ..., -1.4660, -1.4802, -1.4802],
         [-1.4233, -1.4233, -1.4376,  ..., -1.4660, -1.4802, -1.4802],
         ...,
         [-1.3238, -0.8119, -0.5559,  ..., -0.4422, -0.5133, -1.3949],
         [-1.3238, -0.8261, -0.3995,  ..., -0.2857, -0.3142, -1.3380],
         [-1.0821, -0.5133,  0.0698,  ..., -0.3568, -0.1151, -1.2527]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat would be a suggestion for improvement to this breakfast scene? ASSISTANT: A suggestion for improving the breakfast scene could be to add some fresh fruits and/or yogurt to complement the existing meal, consisting of two croissants and a cup of coffee. Including fruit or yogurt would bring more balance to the meal, as it would provide additional nutrients, such as vitamins, minerals, and dietary fiber, that are not as abundant in pastries and coffee alone. This would enhance the overall breakfast experience by providing a more colorful, flavorful, and nutritionally diverse meal for the individual enjoying the breakfast.</s>"], tensor([], size=(0, 428, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (685, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat would be a suggestion for improvement to this breakfast scene? ASSISTANT: A suggestion for improving the breakfast scene could be to add some fresh fruits and/or yogurt to complement the existing meal, consisting of two croissants and a cup of coffee. Including fruit or yogurt would bring more balance to the meal, as it would provide additional nutrients, such as vitamins, minerals, and dietary fiber, that are not as abundant in pastries and coffee alone. This would enhance the overall breakfast experience by providing a more colorful, flavorful, and nutritionally diverse meal for the individual enjoying the breakfast.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat would be a suggestion for improvement to this breakfast scene? ASSISTANT: A suggestion for improving the breakfast scene could be to add some fresh fruits and/or yogurt to complement the existing meal, consisting of two croissants and a cup of coffee. Including fruit or yogurt would bring more balance to the meal, as it would provide additional nutrients, such as vitamins, minerals, and dietary fiber, that are not as abundant in pastries and coffee alone. This would enhance the overall breakfast experience by providing a more colorful, flavorful, and nutritionally diverse meal for the individual enjoying the breakfast.</s>"], [{}], False, 'vqa')]
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000554566.jpg', tensor([[[-1.9638, -1.7240, -1.4158,  ...,  0.0000,  0.0000,  0.0000],
         [-1.9980, -1.8439, -1.6555,  ...,  0.0000,  0.0000,  0.0000],
         [-2.0323, -1.9980, -1.9467,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 0.0056,  0.1939,  0.4166,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0227,  0.1254,  0.2624,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0227,  0.0741,  0.1426,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.8957, -1.6506, -1.3004,  ...,  0.0000,  0.0000,  0.0000],
         [-1.8606, -1.7381, -1.5455,  ...,  0.0000,  0.0000,  0.0000],
         [-1.8256, -1.8431, -1.8256,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 0.3803,  0.5553,  0.7654,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.4678,  0.5378,  0.6429,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.5203,  0.5203,  0.5203,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.4210, -1.2641, -1.0376,  ...,  0.0000,  0.0000,  0.0000],
         [-1.4733, -1.3861, -1.2641,  ...,  0.0000,  0.0000,  0.0000],
         [-1.5604, -1.5430, -1.5256,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 1.1062,  1.2980,  1.5245,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.0888,  1.2631,  1.4374,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.0714,  1.2108,  1.3677,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.5149, -1.4857, -1.4711,  ..., -0.0988, -0.1864, -0.3324],
         [-1.4419, -1.4419, -1.3835,  ..., -0.0405, -0.2010, -0.3178],
         [-1.3105, -1.3251, -1.3981,  ...,  0.0033, -0.1718, -0.2594],
         ...,
         [-0.1572,  0.0325,  0.3829,  ..., -0.5222, -0.5514, -0.5952],
         [-0.5514, -0.5514, -0.0113,  ..., -0.1718, -0.3762, -0.7120],
         [-0.5806, -0.4784,  0.0617,  ..., -0.2302, -0.4346, -0.5660]],

        [[-1.4219, -1.4219, -1.4519,  ...,  0.3490,  0.3940,  0.2139],
         [-1.3769, -1.3619, -1.2718,  ...,  0.3640,  0.3790,  0.2589],
         [-0.4314, -0.3714, -0.2963,  ...,  0.3640,  0.3790,  0.2890],
         ...,
         [-0.0562, -0.0412,  0.3040,  ..., -0.3414, -0.3864, -0.3714],
         [-0.4464, -0.4014, -0.0712,  ...,  0.1389, -0.1613, -0.5065],
         [-0.4164, -0.3264,  0.2139,  ..., -0.0262, -0.1913, -0.0412]],

        [[-1.1105, -1.1247, -1.1247,  ...,  1.0794,  1.3069,  1.2216],
         [-1.0252, -0.9967, -0.9114,  ...,  1.0794,  1.3069,  1.2785],
         [-0.0582,  0.0840,  0.0840,  ...,  1.0367,  1.2785,  1.3069],
         ...,
         [-0.3000, -0.2004, -0.0298,  ...,  0.1977,  0.0840,  0.2120],
         [-0.4848, -0.5133, -0.1578,  ...,  0.6528,  0.3542,  0.0271],
         [-0.4564, -0.3426,  0.0271,  ...,  0.5817,  0.4537,  0.5248]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the black cow standing to the left of the small boy in fence in this image? ASSISTANT: The region corresponding to the black cow standing to the left of the small boy in fence is shown.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the black and white cow in this image? Please give some explanation. ASSISTANT: I've marked the the black and white cow for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is a kid watching cattle in this image? Please give some explanation. ASSISTANT: Sure — this shows the a kid watching cattle.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the little boy in a striped shirt in this image. ASSISTANT: Below you can see the little boy in a striped shirt.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the white cow from the provided image. ASSISTANT: The region corresponding to the white cow is shown.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 691), ['<image>\nCan you segment the black cow standing to the left of the small boy in fence in this image?', '<image>\nWhat is the black and white cow in this image? Please give some explanation.', '<image>\nWhat is a kid watching cattle in this image? Please give some explanation.', '<image>\nPlease highlight the little boy in a striped shirt in this image.', '<image>\nSegment the white cow from the provided image.'], ['black cow standing to the left of the small boy in fence', 'the black and white cow', 'a kid watching cattle', 'little boy in a striped shirt', 'white cow'], [{'black cow standing to the left of the small boy in fence': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the black and white cow': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'a kid watching cattle': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'little boy in a striped shirt': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'white cow': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000211830.jpg', tensor([[[-1.0733, -0.7479, -0.2856,  ...,  2.2147,  2.2318,  2.2489],
         [-1.1932, -1.0390, -0.7993,  ...,  2.2147,  2.2147,  2.2318],
         [-1.3473, -1.3987, -1.4500,  ...,  2.2147,  2.1975,  2.1975],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.7927, -0.4601,  0.0126,  ...,  2.3936,  2.4111,  2.4111],
         [-0.9678, -0.7927, -0.5476,  ...,  2.3936,  2.3936,  2.3936],
         [-1.1954, -1.2129, -1.2479,  ...,  2.3936,  2.3761,  2.3761],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.5495, -0.2358,  0.2173,  ...,  2.6051,  2.6226,  2.6400],
         [-0.7238, -0.5670, -0.3230,  ...,  2.6051,  2.6226,  2.6400],
         [-0.9330, -0.9853, -1.0201,  ...,  2.6051,  2.6226,  2.6226],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.2515,  0.2369,  0.2223,  ...,  1.8865,  1.8281,  1.7552],
         [ 0.2369,  0.2515,  0.2223,  ...,  1.5654,  1.3172,  1.1712],
         [ 0.2515,  0.2369,  0.2515,  ...,  1.0106,  1.0544,  1.1712],
         ...,
         [-1.5587, -1.5441, -1.5441,  ..., -1.4127, -1.3543, -1.3543],
         [-1.5733, -1.5587, -1.5441,  ..., -1.4273, -1.3981, -1.4127],
         [-1.5733, -1.5587, -1.5587,  ..., -1.4127, -1.3981, -1.4127]],

        [[ 0.2589,  0.2439,  0.1989,  ...,  2.0299,  1.9548,  1.8198],
         [ 0.2589,  0.2589,  0.1989,  ...,  1.6547,  1.4145,  1.3095],
         [ 0.2589,  0.2439,  0.2289,  ...,  1.1894,  1.3245,  1.4746],
         ...,
         [-1.4820, -1.4669, -1.4669,  ..., -1.3019, -1.2869, -1.2869],
         [-1.4970, -1.4820, -1.4669,  ..., -1.3169, -1.3319, -1.3469],
         [-1.4970, -1.4820, -1.4820,  ..., -1.3019, -1.3319, -1.3469]],

        [[ 0.3684,  0.3542,  0.3399,  ...,  2.0748,  2.0179,  1.8899],
         [ 0.3684,  0.3684,  0.3399,  ...,  1.6624,  1.4918,  1.3922],
         [ 0.3684,  0.3542,  0.3684,  ...,  1.3211,  1.5060,  1.6340],
         ...,
         [-1.2385, -1.2243, -1.2243,  ..., -1.0110, -0.9825, -0.9825],
         [-1.2527, -1.2385, -1.2243,  ..., -1.0252, -1.0252, -1.0394],
         [-1.2527, -1.2385, -1.2385,  ..., -1.0110, -1.0252, -1.0394]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the bus on left in this image. ASSISTANT: Below you can see the bus on left.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the red bus from the provided image. ASSISTANT: This is the selected red bus.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the left bus in this picture? ASSISTANT: Here's where the left bus appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the left bus in this image. ASSISTANT: Displayed here is the left bus.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the red bus on right in this image? ASSISTANT: Take a look at the red bus on right here.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nPlease highlight the bus on left in this image.', '<image>\nSegment the red bus from the provided image.', '<image>\nCould you isolate the left bus in this picture?', '<image>\nPlease segment the left bus in this image.', '<image>\nCan you segment the red bus on right in this image?'], ['bus on left', 'red bus', 'left bus', 'left bus', 'red bus on right'], [{'bus on left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'red bus': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'left bus': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'left bus': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'red bus on right': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/irbbK8RZH_nfyuLMJVk73g.jpg', tensor([[[1.2043, 1.2043, 1.3242,  ..., 1.4954, 1.4954, 1.4783],
         [1.2557, 1.2728, 1.3927,  ..., 1.4954, 1.5297, 1.5125],
         [1.2043, 1.3242, 1.3584,  ..., 1.5639, 1.5297, 1.5982],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.6057, 1.6057, 1.7283,  ..., 1.7283, 1.7283, 1.7108],
         [1.6583, 1.6758, 1.7983,  ..., 1.7283, 1.7633, 1.7458],
         [1.6057, 1.7283, 1.7633,  ..., 1.7983, 1.7633, 1.8333],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.0474, 2.0474, 2.1694,  ..., 2.1346, 2.1346, 2.1171],
         [2.0997, 2.1171, 2.2391,  ..., 2.1171, 2.1520, 2.1520],
         [2.0474, 2.1694, 2.2043,  ..., 2.1694, 2.1346, 2.2043],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.8573,  1.8573,  1.8865,  ...,  1.8865,  1.8865,  1.8865],
         [ 1.8281,  1.8573,  1.8865,  ...,  1.8865,  1.8865,  1.8865],
         [ 1.8573,  1.8719,  1.8865,  ...,  1.8865,  1.8865,  1.8865],
         ...,
         [-0.9164, -0.9748, -0.9456,  ..., -0.9310, -0.9310, -0.9310],
         [-1.0185, -1.0039, -0.9164,  ..., -0.9310, -0.9164, -0.8580],
         [-1.0039, -0.9602, -0.8872,  ..., -0.9164, -0.9310, -0.8726]],

        [[ 2.0299,  2.0299,  2.0299,  ...,  2.0599,  2.0599,  2.0599],
         [ 1.9998,  2.0149,  2.0299,  ...,  2.0599,  2.0599,  2.0599],
         [ 1.9998,  2.0149,  2.0299,  ...,  2.0599,  2.0599,  2.0599],
         ...,
         [-1.0467, -1.0467, -1.0617,  ..., -0.9117, -0.9117, -0.9417],
         [-1.1218, -1.0918, -1.0317,  ..., -0.8816, -0.8516, -0.8366],
         [-1.0918, -1.0617, -1.0017,  ..., -0.8516, -0.8666, -0.8366]],

        [[ 2.0890,  2.0890,  2.1032,  ...,  2.0606,  2.0606,  2.0606],
         [ 2.0748,  2.0890,  2.1032,  ...,  2.0606,  2.0606,  2.0606],
         [ 2.1032,  2.1032,  2.1032,  ...,  2.0606,  2.0606,  2.0606],
         ...,
         [-0.9683, -0.9967, -0.9967,  ..., -0.7834, -0.7834, -0.7977],
         [-1.0252, -1.0110, -0.9683,  ..., -0.7550, -0.7408, -0.7123],
         [-0.9967, -0.9683, -0.9399,  ..., -0.7408, -0.7550, -0.7266]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the curb shown in this image. ASSISTANT: The curb portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the lane marking - dashed line in this image. ASSISTANT: lane marking - dashed line.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the motorcyclist from the provided image. ASSISTANT: motorcyclist.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is traffic sign (back) in this image? Please elaborate your answer and explain why. ASSISTANT: Sure — this shows the traffic sign (back).</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the traffic sign - direction (front) shown in this image. ASSISTANT: This is the selected traffic sign - direction (front).</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        ...,
        [ 7,  7,  7,  ..., 24, 24, 24],
        [ 7,  7,  7,  ..., 24, 24, 24],
        [ 7,  7,  7,  ..., 24, 24, 24]]), (765, 1024), ['<image>\nDisplay a segmentation mask for the curb shown in this image.', '<image>\nPlease highlight the lane marking - dashed line in this image.', '<image>\nSegment the motorcyclist from the provided image.', '<image>\nWhat is traffic sign (back) in this image? Please elaborate your answer and explain why.', '<image>\nDisplay a segmentation mask for the traffic sign - direction (front) shown in this image.'], ['curb', 'lane marking - dashed line', 'motorcyclist', 'traffic sign (back)', 'traffic sign - direction (front)'], [{'curb': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'lane marking - dashed line': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'motorcyclist': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'traffic sign (back)': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'traffic sign - direction (front)': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000027307.jpg', tensor([[[ 0.5022,  0.5022,  0.4851,  ..., -0.9877, -1.0390, -1.0733],
         [ 0.5022,  0.5022,  0.4851,  ..., -0.8507, -0.9705, -1.0733],
         [ 0.5022,  0.5022,  0.4851,  ..., -0.6794, -0.9020, -1.0562],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.1506,  1.1506,  1.1331,  ..., -0.6702, -0.7577, -0.8102],
         [ 1.1506,  1.1506,  1.1331,  ..., -0.4776, -0.6352, -0.7402],
         [ 1.1506,  1.1506,  1.1331,  ..., -0.2500, -0.4776, -0.6527],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.0648,  2.0648,  2.0474,  ..., -0.2358, -0.4450, -0.6018],
         [ 2.0648,  2.0648,  2.0474,  ..., -0.2358, -0.4798, -0.6715],
         [ 2.0648,  2.0648,  2.0474,  ..., -0.3055, -0.5844, -0.8284],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.4559,  0.4559,  0.4705,  ...,  0.7041,  0.7333,  0.7333],
         [ 0.4705,  0.4705,  0.4705,  ...,  0.7479,  0.7625,  0.7333],
         [ 0.4559,  0.4559,  0.4705,  ...,  0.7625,  0.7917,  0.7771],
         ...,
         [-0.9164, -0.8580, -0.9602,  ...,  0.4851,  0.4267,  0.6311],
         [ 0.4997,  0.3391,  0.4413,  ...,  0.4413,  0.4559,  0.5581],
         [ 0.5143,  0.4705,  0.4851,  ...,  0.5143,  0.6457,  0.5143]],

        [[ 1.0093,  1.0093,  1.0243,  ...,  1.2495,  1.2645,  1.2495],
         [ 1.0243,  1.0243,  1.0243,  ...,  1.2795,  1.2795,  1.2795],
         [ 1.0093,  1.0093,  1.0243,  ...,  1.2795,  1.2795,  1.2945],
         ...,
         [-0.7316, -0.7766, -0.8516,  ...,  0.4841,  0.4090,  0.4540],
         [ 0.5741,  0.4540,  0.5291,  ...,  0.3640,  0.3040,  0.3190],
         [ 0.4991,  0.4691,  0.4991,  ...,  0.4540,  0.5141,  0.3790]],

        [[ 1.6766,  1.6766,  1.6909,  ...,  1.8757,  1.8899,  1.8615],
         [ 1.6909,  1.6909,  1.6909,  ...,  1.8615,  1.9042,  1.8757],
         [ 1.6766,  1.6766,  1.6909,  ...,  1.9184,  1.8615,  1.8899],
         ...,
         [-0.9967, -0.9967, -1.1105,  ..., -0.4564, -0.4848, -0.4422],
         [-0.0156, -0.2004, -0.1009,  ..., -0.5275, -0.4706, -0.3711],
         [-0.3711, -0.5133, -0.4564,  ..., -0.4137, -0.3711, -0.4279]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is 28 in this image? Please give some explanation. ASSISTANT: Here's where the 28 appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the guy on left in this image? ASSISTANT: Here's where the guy on left appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the far left in this image. ASSISTANT: Sure — this shows the far left.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the 28 in this image? ASSISTANT: Here's where the 28 appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the batter in this image? ASSISTANT: You can see the batter in this frame.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (683, 1024), ['<image>\nWhat is 28 in this image? Please give some explanation.', '<image>\nCan you segment the guy on left in this image?', '<image>\nPlease highlight the far left in this image.', '<image>\nCan you segment the 28 in this image?', '<image>\nCan you segment the batter in this image?'], ['28', 'guy on left', 'far left', '28', 'batter'], [{'28': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'guy on left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'far left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'28': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'batter': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/UhlQ8w3D7WiyOF_nMH7dlw.jpg', tensor([[[1.8550, 1.8550, 1.8550,  ..., 0.8447, 0.8104, 0.8276],
         [1.8550, 1.8550, 1.8550,  ..., 0.8104, 0.8276, 0.8447],
         [1.8550, 1.8550, 1.8550,  ..., 0.7933, 0.8447, 0.8447],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.3936, 2.3936, 2.3936,  ..., 1.4832, 1.4482, 1.4657],
         [2.3936, 2.3936, 2.3936,  ..., 1.4482, 1.4657, 1.4832],
         [2.3936, 2.3936, 2.3936,  ..., 1.4307, 1.4832, 1.4832],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.6400, 2.6400, 2.6400,  ..., 1.8905, 1.8557, 1.8731],
         [2.6400, 2.6400, 2.6400,  ..., 1.8557, 1.8731, 1.8905],
         [2.6400, 2.6400, 2.6400,  ..., 1.8383, 1.8905, 1.8905],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.4778,  1.4924,  1.5070,  ...,  0.8938,  0.9376,  0.9522],
         [ 1.4924,  1.4778,  1.4632,  ...,  0.8647,  0.9084,  0.9230],
         [ 1.4632,  1.4778,  1.4632,  ...,  0.8501,  0.8647,  0.8938],
         ...,
         [-1.0915, -1.0623, -1.0477,  ..., -1.0623, -1.0623, -1.0185],
         [-1.1353, -1.1353, -1.0623,  ..., -1.0769, -1.0769, -1.0623],
         [-1.1061, -1.1061, -1.0623,  ..., -1.0769, -1.0769, -1.1061]],

        [[ 1.9848,  1.9998,  2.0149,  ...,  1.3845,  1.3995,  1.4295],
         [ 1.9998,  1.9848,  1.9698,  ...,  1.3845,  1.3845,  1.4295],
         [ 1.9698,  1.9848,  1.9698,  ...,  1.3845,  1.3995,  1.4446],
         ...,
         [-0.7616, -0.7466, -0.7016,  ..., -0.7016, -0.7016, -0.6715],
         [-0.8066, -0.8066, -0.7466,  ..., -0.7166, -0.7166, -0.7016],
         [-0.7766, -0.7766, -0.7466,  ..., -0.7166, -0.7166, -0.7466]],

        [[ 2.1317,  2.1459,  2.1317,  ...,  1.6340,  1.6482,  1.6482],
         [ 2.1459,  2.1317,  2.1459,  ...,  1.6198,  1.6482,  1.6482],
         [ 2.1459,  2.1459,  2.1459,  ...,  1.6482,  1.6624,  1.6624],
         ...,
         [-0.5417, -0.5275, -0.4848,  ..., -0.3853, -0.3711, -0.3853],
         [-0.5844, -0.5986, -0.5275,  ..., -0.3995, -0.3995, -0.3853],
         [-0.5559, -0.5559, -0.5275,  ..., -0.4137, -0.3995, -0.4279]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the signage - advertisement in this picture? ASSISTANT: Sure — this shows the signage - advertisement.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is static in this image? Please give some explanation. ASSISTANT: Displayed here is the static.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the terrain in this picture? ASSISTANT: Displayed here is the terrain.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the traffic sign (front) in this image? ASSISTANT: This view focuses on the traffic sign (front).</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the curb in this picture? ASSISTANT: I've marked the curb for you.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        ...,
        [21, 21, 21,  ..., 21, 21, 21],
        [21, 21, 21,  ..., 21, 21, 21],
        [21, 21, 21,  ..., 21, 21, 21]]), (765, 1024), ['<image>\nCould you isolate the signage - advertisement in this picture?', '<image>\nWhat is static in this image? Please give some explanation.', '<image>\nCould you isolate the terrain in this picture?', '<image>\nCan you segment the traffic sign (front) in this image?', '<image>\nCould you identify and segment out the curb in this picture?'], ['signage - advertisement', 'static', 'terrain', 'traffic sign (front)', 'curb'], [{'signage - advertisement': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'static': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'terrain': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'traffic sign (front)': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'curb': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000513748.jpg', tensor([[[-1.5357, -1.5185, -1.4843,  ..., -1.7412, -1.8268, -1.8953],
         [-1.5357, -1.5014, -1.4672,  ..., -1.8268, -1.8610, -1.8953],
         [-1.5185, -1.4843, -1.4329,  ..., -1.9467, -1.9124, -1.8782],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.4580, -1.4405, -1.4055,  ..., -1.6506, -1.7556, -1.8256],
         [-1.4580, -1.4230, -1.3880,  ..., -1.7206, -1.7731, -1.8081],
         [-1.4405, -1.4055, -1.3529,  ..., -1.8256, -1.8081, -1.7906],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.3164, -1.2990, -1.2641,  ..., -1.4210, -1.4733, -1.5081],
         [-1.3164, -1.2816, -1.2467,  ..., -1.5430, -1.5430, -1.5256],
         [-1.2990, -1.2641, -1.2119,  ..., -1.6824, -1.6127, -1.5604],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.5003, -1.2667, -1.2521,  ..., -1.6755, -1.6901, -1.6755],
         [-1.5587, -1.2229, -1.0039,  ..., -1.6755, -1.6609, -1.6755],
         [-1.6317, -1.4419, -1.4857,  ..., -1.6609, -1.6609, -1.6463],
         ...,
         [ 0.5289,  0.4997,  0.5727,  ...,  1.2442,  1.2588,  1.2004],
         [ 0.9084,  0.7041,  0.7771,  ...,  1.1566,  1.1858,  1.2296],
         [ 0.5435,  0.6895,  0.6895,  ...,  1.2296,  1.2150,  1.2442]],

        [[-1.4369, -1.1968, -1.2418,  ..., -1.6320, -1.6470, -1.6320],
         [-1.4970, -1.1518, -0.9717,  ..., -1.6320, -1.6170, -1.6320],
         [-1.5570, -1.3619, -1.4219,  ..., -1.6170, -1.6170, -1.6020],
         ...,
         [ 0.7992,  0.7542,  0.8142,  ...,  1.4446,  1.5196,  1.4295],
         [ 1.1444,  0.9193,  0.9793,  ...,  1.3545,  1.4295,  1.4596],
         [ 0.7542,  0.8893,  0.9193,  ...,  1.4295,  1.4145,  1.4446]],

        [[-1.2954, -1.1105, -1.0963,  ..., -1.3665, -1.3807, -1.3665],
         [-1.3522, -1.0394, -0.8261,  ..., -1.3665, -1.3522, -1.3665],
         [-1.3807, -1.1958, -1.2385,  ..., -1.3522, -1.3522, -1.3380],
         ...,
         [ 0.9088,  0.9088,  0.9088,  ...,  1.6055,  1.6340,  1.5629],
         [ 1.2785,  1.1078,  1.1221,  ...,  1.5202,  1.5629,  1.6055],
         [ 0.8803,  1.0510,  1.0367,  ...,  1.5771,  1.5629,  1.5913]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nHow many people are visible surfing in the image? ASSISTANT: There are many people visible in the image, riding surfboards in the water.</s>USER: Are the surfers waiting in line to surf the waves, or are they all surfing at the same time? ASSISTANT: The surfers are all surfing at the same time, rather than waiting in line to surf the waves.</s>USER: Are the waves in the image natural or simulated in a wave pool? ASSISTANT: The exact nature of the waves cannot be confirmed with certainty from the image, as they could potentially be either natural waves in the ocean or simulated waves in a wave pool.</s>USER: What is the size of the waves that the surfers are riding? ASSISTANT: The surfers are riding small waves in the water.</s>USER: What is a simulated wave pool and how does it work? ASSISTANT: A simulated wave pool is a specially designed swimming pool that generates artificial waves to mimic the conditions experienced when surfing in the ocean. Wave pools are created for recreational purposes and for learning how to surf, providing a controlled environment for surfers to practice their skills without the risks and challenges associated with ocean surfing.\n\nWave pools typically use one of several wave-generating technologies to create waves. These include pneumatic systems, hydraulic systems, and paddles systems. Some of the more advanced wave pools use a combination of these technologies to produce customizable waves that can cater to the needs and skill levels of different surfers.\n\nPneumatic systems utilize air pressure to move water and generate waves. This method involves compressing air in submerged chambers, which then get rapidly released into the pool, displacing water and creating a wave.\n\nHydraulic systems, on the other hand, use underwater pistons or plungers that move vertically to generate waves. These pistons push water upward, creating a wave that travels across the pool.\n\nPaddle systems use paddle-like structures positioned around the pool's perimeter, which move back and forth to create waves by pushing and pulling water in a specific direction.\n\nOverall, simulated wave pools offer a controlled environment for surfers of all skill levels to practice their techniques and have fun, without the unpredictability and potential hazards found in natural ocean conditions.</s>"], tensor([], size=(0, 424, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (678, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nHow many people are visible surfing in the image? ASSISTANT: There are many people visible in the image, riding surfboards in the water.</s>USER: Are the surfers waiting in line to surf the waves, or are they all surfing at the same time? ASSISTANT: The surfers are all surfing at the same time, rather than waiting in line to surf the waves.</s>USER: Are the waves in the image natural or simulated in a wave pool? ASSISTANT: The exact nature of the waves cannot be confirmed with certainty from the image, as they could potentially be either natural waves in the ocean or simulated waves in a wave pool.</s>USER: What is the size of the waves that the surfers are riding? ASSISTANT: The surfers are riding small waves in the water.</s>USER: What is a simulated wave pool and how does it work? ASSISTANT: A simulated wave pool is a specially designed swimming pool that generates artificial waves to mimic the conditions experienced when surfing in the ocean. Wave pools are created for recreational purposes and for learning how to surf, providing a controlled environment for surfers to practice their skills without the risks and challenges associated with ocean surfing.\n\nWave pools typically use one of several wave-generating technologies to create waves. These include pneumatic systems, hydraulic systems, and paddles systems. Some of the more advanced wave pools use a combination of these technologies to produce customizable waves that can cater to the needs and skill levels of different surfers.\n\nPneumatic systems utilize air pressure to move water and generate waves. This method involves compressing air in submerged chambers, which then get rapidly released into the pool, displacing water and creating a wave.\n\nHydraulic systems, on the other hand, use underwater pistons or plungers that move vertically to generate waves. These pistons push water upward, creating a wave that travels across the pool.\n\nPaddle systems use paddle-like structures positioned around the pool's perimeter, which move back and forth to create waves by pushing and pulling water in a specific direction.\n\nOverall, simulated wave pools offer a controlled environment for surfers of all skill levels to practice their techniques and have fun, without the unpredictability and potential hazards found in natural ocean conditions.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nHow many people are visible surfing in the image? ASSISTANT: There are many people visible in the image, riding surfboards in the water.</s>USER: Are the surfers waiting in line to surf the waves, or are they all surfing at the same time? ASSISTANT: The surfers are all surfing at the same time, rather than waiting in line to surf the waves.</s>USER: Are the waves in the image natural or simulated in a wave pool? ASSISTANT: The exact nature of the waves cannot be confirmed with certainty from the image, as they could potentially be either natural waves in the ocean or simulated waves in a wave pool.</s>USER: What is the size of the waves that the surfers are riding? ASSISTANT: The surfers are riding small waves in the water.</s>USER: What is a simulated wave pool and how does it work? ASSISTANT: A simulated wave pool is a specially designed swimming pool that generates artificial waves to mimic the conditions experienced when surfing in the ocean. Wave pools are created for recreational purposes and for learning how to surf, providing a controlled environment for surfers to practice their skills without the risks and challenges associated with ocean surfing.\n\nWave pools typically use one of several wave-generating technologies to create waves. These include pneumatic systems, hydraulic systems, and paddles systems. Some of the more advanced wave pools use a combination of these technologies to produce customizable waves that can cater to the needs and skill levels of different surfers.\n\nPneumatic systems utilize air pressure to move water and generate waves. This method involves compressing air in submerged chambers, which then get rapidly released into the pool, displacing water and creating a wave.\n\nHydraulic systems, on the other hand, use underwater pistons or plungers that move vertically to generate waves. These pistons push water upward, creating a wave that travels across the pool.\n\nPaddle systems use paddle-like structures positioned around the pool's perimeter, which move back and forth to create waves by pushing and pulling water in a specific direction.\n\nOverall, simulated wave pools offer a controlled environment for surfers of all skill levels to practice their techniques and have fun, without the unpredictability and potential hazards found in natural ocean conditions.</s>"], [{}], False, 'vqa')]
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/saiapr_tc-12/02/images/2017.jpg', tensor([[[0.8789, 0.6906, 0.2796,  ..., 1.0844, 1.0844, 1.0844],
         [0.8618, 0.6906, 0.3481,  ..., 1.1015, 1.1015, 1.1015],
         [0.7933, 0.7077, 0.5022,  ..., 1.1358, 1.1529, 1.1529],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.2206, 1.0280, 0.6078,  ..., 1.8683, 1.8683, 1.8683],
         [1.2031, 1.0280, 0.6779,  ..., 1.8859, 1.8859, 1.8859],
         [1.1331, 1.0455, 0.8354,  ..., 1.9209, 1.9384, 1.9384],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.5420, 1.3502, 0.9319,  ..., 2.2914, 2.2914, 2.2914],
         [1.5245, 1.3502, 1.0017,  ..., 2.3088, 2.3088, 2.3088],
         [1.4548, 1.3677, 1.1585,  ..., 2.3437, 2.3611, 2.3611],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.8573,  1.8865,  1.8719,  ..., -1.2375, -0.9456, -1.0185],
         [ 1.7698,  1.8135,  1.7552,  ..., -0.8872, -0.3324, -0.7996],
         [ 1.2004,  1.2004,  1.0982,  ..., -0.3032,  0.4121, -0.4930],
         ...,
         [-1.6755, -1.7777, -1.7777,  ..., -1.7193, -1.7923, -1.2813],
         [-1.6755, -1.7339, -1.7923,  ..., -1.7631, -1.7339, -1.4565],
         [-1.6609, -1.6901, -1.7923,  ..., -1.7485, -1.7923, -1.6317]],

        [[ 2.0749,  2.0749,  2.0749,  ..., -0.7616, -0.5365, -0.5815],
         [ 2.0149,  2.0299,  1.9848,  ..., -0.2963,  0.2589, -0.2063],
         [ 1.4896,  1.4896,  1.3995,  ...,  0.5291,  1.2344,  0.2740],
         ...,
         [-1.5870, -1.5720, -1.5270,  ..., -1.6621, -1.7371, -1.2268],
         [-1.5720, -1.5120, -1.5570,  ..., -1.7221, -1.6921, -1.4219],
         [-1.5270, -1.4669, -1.5420,  ..., -1.7071, -1.7521, -1.5870]],

        [[ 2.0464,  2.0464,  2.1032,  ..., -0.2004,  0.0555,  0.0413],
         [ 2.0321,  2.0321,  2.0464,  ...,  0.2262,  0.7666,  0.3257],
         [ 1.5629,  1.5487,  1.5202,  ...,  0.9514,  1.6340,  0.7523],
         ...,
         [-1.1674, -1.0678, -0.9825,  ..., -1.2527, -1.3949, -0.9399],
         [-1.1532, -0.9967, -1.0110,  ..., -1.3096, -1.3380, -1.1389],
         [-1.1105, -0.9399, -0.9967,  ..., -1.2811, -1.3949, -1.2954]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is window upper right in this image? Please elaborate your answer and explain why. ASSISTANT: You can see the window upper right in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is lower middle boy 's face in this image? Please give some explanation. ASSISTANT: This view focuses on the lower middle boy 's face.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the boy in upper left in this picture? ASSISTANT: The boy in upper left is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the face of kid at the top center in this picture? ASSISTANT: face of kid at the top center.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the face of boy on left shown in this image. ASSISTANT: You can see the face of boy on left in this frame.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nWhat is window upper right in this image? Please elaborate your answer and explain why.', "<image>\nWhat is lower middle boy 's face in this image? Please give some explanation.", '<image>\nCould you identify and segment out the boy in upper left in this picture?', '<image>\nCould you isolate the face of kid at the top center in this picture?', '<image>\nDisplay a segmentation mask for the face of boy on left shown in this image.'], ['window upper right', "lower middle boy 's face", 'boy in upper left', 'face of kid at the top center', 'face of boy on left'], [{'window upper right': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {"lower middle boy 's face": tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'boy in upper left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'face of kid at the top center': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'face of boy on left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000073690.jpg', tensor([[[-0.3541, -0.3541, -0.3369,  ..., -1.3302, -1.3473, -1.3473],
         [-0.3369, -0.3369, -0.3369,  ..., -1.3473, -1.3473, -1.3473],
         [-0.3198, -0.3198, -0.3198,  ..., -1.3644, -1.3644, -1.3644],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.1800, -0.1800, -0.1625,  ..., -1.3179, -1.3354, -1.3354],
         [-0.1625, -0.1625, -0.1625,  ..., -1.3354, -1.3354, -1.3354],
         [-0.1450, -0.1450, -0.1450,  ..., -1.3529, -1.3529, -1.3529],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.2532, -0.2532, -0.2358,  ..., -1.1596, -1.1770, -1.1770],
         [-0.2358, -0.2358, -0.2358,  ..., -1.1770, -1.1770, -1.1770],
         [-0.2184, -0.2184, -0.2184,  ..., -1.1944, -1.1944, -1.1944],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.4121,  0.4121,  0.4267,  ..., -0.4492, -0.4638, -0.4638],
         [ 0.4267,  0.4413,  0.4413,  ..., -0.4784, -0.4784, -0.4784],
         [ 0.4267,  0.4267,  0.4267,  ..., -0.4784, -0.4930, -0.4930],
         ...,
         [-0.1134, -1.1645, -1.3835,  ..., -1.7047, -1.7047, -1.7193],
         [-1.1207, -1.3105, -1.6755,  ..., -1.6901, -1.7193, -1.7047],
         [-1.2229, -1.6609, -1.6171,  ..., -1.6901, -1.7047, -1.7047]],

        [[ 0.4841,  0.4841,  0.4991,  ..., -0.5215, -0.4914, -0.5215],
         [ 0.4841,  0.4841,  0.4991,  ..., -0.5365, -0.5215, -0.5365],
         [ 0.4691,  0.4691,  0.4691,  ..., -0.5365, -0.5365, -0.5515],
         ...,
         [-0.1463, -1.2118, -1.2869,  ..., -1.6771, -1.6771, -1.6921],
         [-1.1218, -1.2718, -1.5720,  ..., -1.6621, -1.6921, -1.6771],
         [-1.1668, -1.6170, -1.6470,  ..., -1.6621, -1.6771, -1.6771]],

        [[ 0.1835,  0.1977,  0.1977,  ..., -0.3426, -0.3711, -0.3853],
         [ 0.1977,  0.2120,  0.2120,  ..., -0.3995, -0.3995, -0.3995],
         [ 0.2120,  0.2120,  0.2120,  ..., -0.4137, -0.3995, -0.4137],
         ...,
         [-0.1435, -1.1247, -1.2100,  ..., -1.4660, -1.4660, -1.4802],
         [-1.0394, -1.1816, -1.3665,  ..., -1.4518, -1.4802, -1.4660],
         [-1.0963, -1.4233, -1.4091,  ..., -1.4518, -1.4660, -1.4660]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the a man in an orange shirt playing a video game shown in this image. ASSISTANT: Take a look at the a man in an orange shirt playing a video game here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the guy in orange shirt holding a wii remote in this picture. ASSISTANT: The guy in orange shirt holding a wii remote is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the a girl in a green halter top plays wii in this picture. ASSISTANT: The region corresponding to the a girl in a green halter top plays wii is shown.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the the woman wearing the green tank top and blue jeans from the provided image. ASSISTANT: The the woman wearing the green tank top and blue jeans portion is displayed below.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (680, 1024), ['<image>\nDisplay a segmentation mask for the a man in an orange shirt playing a video game shown in this image.', '<image>\nPlease point out the guy in orange shirt holding a wii remote in this picture.', '<image>\nPlease point out the a girl in a green halter top plays wii in this picture.', '<image>\nSegment the the woman wearing the green tank top and blue jeans from the provided image.'], ['a man in an orange shirt playing a video game', 'guy in orange shirt holding a wii remote', 'a girl in a green halter top plays wii', 'the woman wearing the green tank top and blue jeans'], [{'a man in an orange shirt playing a video game': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'guy in orange shirt holding a wii remote': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'a girl in a green halter top plays wii': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the woman wearing the green tank top and blue jeans': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000467878.jpg', tensor([[[-0.3883, -0.3541, -0.2856,  ...,  0.6049,  0.5878,  0.5878],
         [-0.3712, -0.3369, -0.2856,  ...,  0.5707,  0.5707,  0.5878],
         [-0.3369, -0.3027, -0.2684,  ...,  0.5193,  0.5536,  0.5707],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.2325, -0.1975, -0.1275,  ...,  0.7129,  0.6954,  0.6954],
         [-0.2325, -0.1975, -0.1275,  ...,  0.6779,  0.6779,  0.6954],
         [-0.2150, -0.1800, -0.1625,  ...,  0.6429,  0.6604,  0.6779],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.0615, -0.0441,  0.0082,  ...,  1.0191,  1.0017,  1.0017],
         [-0.0092,  0.0082,  0.0431,  ...,  0.9842,  0.9842,  1.0017],
         [ 0.0431,  0.0605,  0.0779,  ...,  0.9494,  0.9668,  0.9842],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.1134, -0.1426, -0.0842,  ...,  0.7333,  0.7333,  0.7479],
         [-0.2740, -0.4054, -0.3324,  ...,  0.6895,  0.7625,  0.7917],
         [-0.1864, -0.4200, -0.4492,  ...,  0.6749,  0.7479,  0.8647],
         ...,
         [-0.0842, -0.0550, -0.0988,  ...,  0.4559,  0.4413,  0.4121],
         [-0.2156, -0.2010, -0.2156,  ...,  0.4851,  0.6165,  0.5289],
         [-0.3032, -0.2886, -0.2302,  ...,  0.4413,  0.6603,  0.5289]],

        [[ 0.0338, -0.0412, -0.0112,  ...,  0.7992,  0.8442,  0.9043],
         [-0.1313, -0.3264, -0.2663,  ...,  0.7992,  0.8593,  0.9343],
         [-0.1613, -0.4164, -0.4014,  ...,  0.7992,  0.8593,  1.0093],
         ...,
         [-0.5965, -0.5815, -0.6115,  ...,  0.5891,  0.5591,  0.5441],
         [-0.5665, -0.5665, -0.5665,  ...,  0.6041,  0.7392,  0.6792],
         [-0.4464, -0.4464, -0.3864,  ...,  0.5741,  0.7842,  0.6642]],

        [[ 0.1124,  0.0698,  0.1124,  ...,  0.8661,  0.9230,  0.9799],
         [-0.0298, -0.2004, -0.0724,  ...,  0.8945,  0.9372,  0.9799],
         [ 0.0555, -0.1151, -0.0867,  ...,  0.8945,  1.0083,  1.1221],
         ...,
         [-0.6270, -0.6128, -0.6270,  ...,  0.6386,  0.6386,  0.6528],
         [-0.5986, -0.5844, -0.5844,  ...,  0.6528,  0.8092,  0.7666],
         [-0.4706, -0.4706, -0.3853,  ...,  0.6386,  0.9088,  0.7808]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat type of food item is being held in the person's hand? ASSISTANT: A grilled hot dog in a bun is being held in the person's hand.</s>USER: Is the person holding the grilled hot dog directly or with something else such as a napkin? ASSISTANT: The person is holding the grilled hot dog with a napkin.</s>USER: How does the hot dog appear to be cooked? ASSISTANT: The hot dog appears to be grilled, which likely gives it a slightly smoky flavor and appealing grill marks on the surface.</s>USER: What could be some popular toppings for the grilled hot dog being held by the person? ASSISTANT: Popular toppings for grilled hot dogs can vary depending on personal preferences and regional influences. Some common toppings include:\n\n1. Ketchup and mustard: These are classic condiments that add a combination of sweet, tangy, and savory flavors to the hot dog.\n2. Relish: This is a pickled cucumber spread, which adds a sweet and tangy crunch to the hot dog.\n3. Onions: Finely chopped raw or grilled onions can add a sharp, pungent taste or a sweet, caramelized flavor depending on how they are prepared.\n4. Sauerkraut: This fermented cabbage topping adds a tangy and slightly sour flavor, which complements the savory taste of the hot dog.\n5. Chili: A warm, meaty chili can turn a hot dog into a hearty, satisfying meal with a kick of spice.\n\nOther possible toppings include cheese, jalapeños, mayonnaise, coleslaw, or even specialty sauces that are unique to specific regions or restaurants. The best topping combination ultimately depends on the individual's taste preferences and any cultural traditions they might wish to embrace in their hot dog preparation.</s>"], tensor([], size=(0, 480, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat type of food item is being held in the person's hand? ASSISTANT: A grilled hot dog in a bun is being held in the person's hand.</s>USER: Is the person holding the grilled hot dog directly or with something else such as a napkin? ASSISTANT: The person is holding the grilled hot dog with a napkin.</s>USER: How does the hot dog appear to be cooked? ASSISTANT: The hot dog appears to be grilled, which likely gives it a slightly smoky flavor and appealing grill marks on the surface.</s>USER: What could be some popular toppings for the grilled hot dog being held by the person? ASSISTANT: Popular toppings for grilled hot dogs can vary depending on personal preferences and regional influences. Some common toppings include:\n\n1. Ketchup and mustard: These are classic condiments that add a combination of sweet, tangy, and savory flavors to the hot dog.\n2. Relish: This is a pickled cucumber spread, which adds a sweet and tangy crunch to the hot dog.\n3. Onions: Finely chopped raw or grilled onions can add a sharp, pungent taste or a sweet, caramelized flavor depending on how they are prepared.\n4. Sauerkraut: This fermented cabbage topping adds a tangy and slightly sour flavor, which complements the savory taste of the hot dog.\n5. Chili: A warm, meaty chili can turn a hot dog into a hearty, satisfying meal with a kick of spice.\n\nOther possible toppings include cheese, jalapeños, mayonnaise, coleslaw, or even specialty sauces that are unique to specific regions or restaurants. The best topping combination ultimately depends on the individual's taste preferences and any cultural traditions they might wish to embrace in their hot dog preparation.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat type of food item is being held in the person's hand? ASSISTANT: A grilled hot dog in a bun is being held in the person's hand.</s>USER: Is the person holding the grilled hot dog directly or with something else such as a napkin? ASSISTANT: The person is holding the grilled hot dog with a napkin.</s>USER: How does the hot dog appear to be cooked? ASSISTANT: The hot dog appears to be grilled, which likely gives it a slightly smoky flavor and appealing grill marks on the surface.</s>USER: What could be some popular toppings for the grilled hot dog being held by the person? ASSISTANT: Popular toppings for grilled hot dogs can vary depending on personal preferences and regional influences. Some common toppings include:\n\n1. Ketchup and mustard: These are classic condiments that add a combination of sweet, tangy, and savory flavors to the hot dog.\n2. Relish: This is a pickled cucumber spread, which adds a sweet and tangy crunch to the hot dog.\n3. Onions: Finely chopped raw or grilled onions can add a sharp, pungent taste or a sweet, caramelized flavor depending on how they are prepared.\n4. Sauerkraut: This fermented cabbage topping adds a tangy and slightly sour flavor, which complements the savory taste of the hot dog.\n5. Chili: A warm, meaty chili can turn a hot dog into a hearty, satisfying meal with a kick of spice.\n\nOther possible toppings include cheese, jalapeños, mayonnaise, coleslaw, or even specialty sauces that are unique to specific regions or restaurants. The best topping combination ultimately depends on the individual's taste preferences and any cultural traditions they might wish to embrace in their hot dog preparation.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000143717.jpg', tensor([[[-0.4226, -0.4054, -0.3712,  ..., -0.9020, -0.9020, -0.9020],
         [-0.4054, -0.3883, -0.3712,  ..., -0.9020, -0.9020, -0.9020],
         [-0.3883, -0.3712, -0.3541,  ..., -0.8849, -0.8849, -0.8849],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.4678,  0.4853,  0.5203,  ...,  0.1176,  0.1176,  0.1176],
         [ 0.4853,  0.5028,  0.5378,  ...,  0.1176,  0.1176,  0.1176],
         [ 0.5203,  0.5378,  0.5553,  ...,  0.1352,  0.1352,  0.1352],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.3677,  1.3851,  1.4200,  ...,  1.1585,  1.1585,  1.1585],
         [ 1.3851,  1.4025,  1.4374,  ...,  1.1585,  1.1585,  1.1585],
         [ 1.4200,  1.4374,  1.4548,  ...,  1.1759,  1.1759,  1.1759],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.0550, -0.0405, -0.0113,  ..., -0.4784, -0.4784, -0.4930],
         [-0.0113,  0.0033,  0.0033,  ..., -0.4492, -0.4638, -0.4638],
         [ 0.0325,  0.0179,  0.0179,  ..., -0.4054, -0.4346, -0.4492],
         ...,
         [-1.7339, -1.7193, -1.6901,  ..., -1.7777, -1.7777, -1.7923],
         [-1.7047, -1.6901, -1.6901,  ..., -1.7777, -1.7777, -1.7631],
         [-1.7339, -1.7339, -1.7339,  ..., -1.7777, -1.7777, -1.7923]],

        [[ 0.6642,  0.6491,  0.6642,  ...,  0.3640,  0.3790,  0.3940],
         [ 0.6642,  0.6642,  0.6792,  ...,  0.3940,  0.4090,  0.4090],
         [ 0.6942,  0.6792,  0.6792,  ...,  0.4240,  0.4090,  0.4240],
         ...,
         [-1.6470, -1.6170, -1.6020,  ..., -1.1818, -1.2268, -1.2268],
         [-1.6170, -1.5870, -1.6170,  ..., -1.1518, -1.1518, -1.1218],
         [-1.6320, -1.6320, -1.6621,  ..., -1.1968, -1.2118, -1.2268]],

        [[ 1.3638,  1.3496,  1.3638,  ...,  1.1932,  1.1932,  1.1789],
         [ 1.3638,  1.3638,  1.3780,  ...,  1.2074,  1.1932,  1.1932],
         [ 1.3922,  1.3780,  1.3780,  ...,  1.2216,  1.2074,  1.2074],
         ...,
         [-1.3949, -1.3665, -1.3522,  ..., -0.5701, -0.6270, -0.6270],
         [-1.3522, -1.3380, -1.3665,  ..., -0.5417, -0.5559, -0.5133],
         [-1.3949, -1.3949, -1.4233,  ..., -0.6697, -0.6412, -0.6412]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat activities can visitors enjoy at this resort besides swimming in the pool? ASSISTANT: At this resort, visitors can find several activities to enjoy besides swimming in the large blue pool. They can relax in the numerous lawn chairs placed around the pool area, which are covered by umbrellas to provide shade from the sun. The resort's proximity to the ocean also offers an opportunity to enjoy beach activities or take a walk along the shoreline. Additionally, visitors can interact with other guests, as there are a number of people present in the image. Overall, the resort offers a relaxing and leisurely environment for guests to unwind and partake in a variety of activities.</s>"], tensor([], size=(0, 427, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (683, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat activities can visitors enjoy at this resort besides swimming in the pool? ASSISTANT: At this resort, visitors can find several activities to enjoy besides swimming in the large blue pool. They can relax in the numerous lawn chairs placed around the pool area, which are covered by umbrellas to provide shade from the sun. The resort's proximity to the ocean also offers an opportunity to enjoy beach activities or take a walk along the shoreline. Additionally, visitors can interact with other guests, as there are a number of people present in the image. Overall, the resort offers a relaxing and leisurely environment for guests to unwind and partake in a variety of activities.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat activities can visitors enjoy at this resort besides swimming in the pool? ASSISTANT: At this resort, visitors can find several activities to enjoy besides swimming in the large blue pool. They can relax in the numerous lawn chairs placed around the pool area, which are covered by umbrellas to provide shade from the sun. The resort's proximity to the ocean also offers an opportunity to enjoy beach activities or take a walk along the shoreline. Additionally, visitors can interact with other guests, as there are a number of people present in the image. Overall, the resort offers a relaxing and leisurely environment for guests to unwind and partake in a variety of activities.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/ade20k/images/training/ADE_train_00008645.jpg', tensor([[[1.6838, 0.9817, 0.7419,  ..., 2.2489, 2.2489, 2.2489],
         [1.7009, 0.9988, 0.7591,  ..., 2.2489, 2.2489, 2.2489],
         [1.7180, 0.9988, 0.7591,  ..., 2.2489, 2.2489, 2.2489],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.2031, 0.4853, 0.2052,  ..., 2.4286, 2.4286, 2.4286],
         [1.2206, 0.5028, 0.2227,  ..., 2.4286, 2.4286, 2.4286],
         [1.2381, 0.5028, 0.2227,  ..., 2.4286, 2.4286, 2.4286],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.1411, 0.4091, 0.1128,  ..., 2.6400, 2.6400, 2.6400],
         [1.1585, 0.4265, 0.1302,  ..., 2.6400, 2.6400, 2.6400],
         [1.1759, 0.4265, 0.1302,  ..., 2.6400, 2.6400, 2.6400],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.7260,  1.7260,  1.3026,  ..., -0.2740, -0.2594, -0.2594],
         [ 1.5362,  1.6384,  1.6676,  ..., -0.2594, -0.2448, -0.1572],
         [ 1.3756,  1.4340,  1.4194,  ..., -0.2594, -0.2156, -0.1572],
         ...,
         [ 0.9668,  0.9668,  1.0106,  ...,  1.9011,  1.9157,  1.9157],
         [ 1.9011,  1.8573,  1.8573,  ...,  1.9011,  1.8865,  1.8865],
         [ 1.8865,  1.8573,  1.8573,  ...,  1.9303,  1.9157,  1.9157]],

        [[ 1.5946,  1.6096,  1.2044,  ..., -0.3264, -0.3564, -0.3564],
         [ 1.4295,  1.5496,  1.5946,  ..., -0.3864, -0.3714, -0.2663],
         [ 1.2945,  1.3695,  1.3395,  ..., -0.3264, -0.2513, -0.1613],
         ...,
         [ 0.9943,  1.0243,  1.0694,  ...,  2.0299,  2.0449,  2.0599],
         [ 2.0599,  2.0299,  2.0149,  ...,  2.0599,  2.0599,  2.0449],
         [ 2.0449,  2.0149,  1.9998,  ...,  2.0749,  2.0599,  2.0449]],

        [[ 0.6244,  0.7097,  0.4964,  ..., -0.0156, -0.0724, -0.0867],
         [ 0.4395,  0.6244,  0.8092,  ..., -0.0298, -0.0298,  0.0413],
         [ 0.3257,  0.4395,  0.4821,  ...,  0.0129,  0.0555,  0.1124],
         ...,
         [ 0.8661,  0.8803,  0.9372,  ...,  2.0464,  2.0606,  2.0748],
         [ 1.8331,  1.8046,  1.8331,  ...,  2.1175,  2.1175,  2.1175],
         [ 1.6482,  1.6198,  1.6482,  ...,  2.0890,  2.0890,  2.0748]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the box in this image. ASSISTANT: The box is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is shelf in this image? Please give some explanation. ASSISTANT: Take a look at the shelf here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is case in this image? Please give some explanation. ASSISTANT: Result: case.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the painting in this picture. ASSISTANT: Result: painting.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the ceiling shown in this image. ASSISTANT: Displayed here is the ceiling.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[ 24,  24,  24,  ...,   0, 255, 255],
        [ 24,  24,  24,  ...,   0, 255, 255],
        [ 24,  24,  24,  ...,   0, 255, 255],
        ...,
        [255, 255, 255,  ...,   3,   3, 255],
        [255, 255, 255,  ...,   3,   3, 255],
        [255, 255, 255,  ..., 255, 255, 255]]), (683, 1024), ['<image>\nPlease highlight the box in this image.', '<image>\nWhat is shelf in this image? Please give some explanation.', '<image>\nWhat is case in this image? Please give some explanation.', '<image>\nPlease point out the painting in this picture.', '<image>\nDisplay a segmentation mask for the ceiling shown in this image.'], ['box', 'shelf', 'case', 'painting', 'ceiling'], [{'box': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'shelf': tensor([[ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'case': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'painting': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'ceiling': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000528466.jpg', tensor([[[-0.9705, -0.9705, -0.9877,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9534, -0.9705, -1.0048,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9363, -0.9705, -1.0219,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.8849, -0.8678, -0.8335,  ...,  0.0000,  0.0000,  0.0000],
         [-0.8507, -0.8164, -0.7822,  ...,  0.0000,  0.0000,  0.0000],
         [-0.8164, -0.7822, -0.7479,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.0028, -1.0028, -1.0028,  ...,  0.0000,  0.0000,  0.0000],
         [-1.0028, -1.0028, -1.0203,  ...,  0.0000,  0.0000,  0.0000],
         [-1.0028, -1.0203, -1.0553,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.7227, -0.7052, -0.6877,  ...,  0.0000,  0.0000,  0.0000],
         [-0.7052, -0.7052, -0.7052,  ...,  0.0000,  0.0000,  0.0000],
         [-0.7052, -0.7052, -0.7052,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.1073, -1.1770, -1.2467,  ...,  0.0000,  0.0000,  0.0000],
         [-1.1596, -1.2119, -1.2293,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2467, -1.2467, -1.1770,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.8458, -0.8284, -0.7936,  ...,  0.0000,  0.0000,  0.0000],
         [-0.8981, -0.7936, -0.6367,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9330, -0.7587, -0.5147,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.7558, -0.7412, -0.7996,  ..., -1.1937, -1.1645, -1.1353],
         [-0.7412, -0.7266, -0.7704,  ..., -1.2229, -1.1937, -1.1499],
         [-0.7120, -0.7704, -0.7412,  ..., -1.1499, -1.1791, -1.2229],
         ...,
         [-0.7996, -0.8142, -0.7558,  ...,  1.4778,  1.5362,  1.4778],
         [-0.7558, -0.8288, -0.7704,  ...,  1.5362,  1.5216,  1.5216],
         [-0.7704, -0.8434, -0.7996,  ...,  1.6092,  1.4924,  1.4486]],

        [[-0.6565, -0.6565, -0.7016,  ..., -1.1668, -1.1218, -1.0918],
         [-0.6415, -0.6415, -0.6565,  ..., -1.1968, -1.1518, -1.0918],
         [-0.6265, -0.6865, -0.6715,  ..., -1.1368, -1.1218, -1.1218],
         ...,
         [-0.7016, -0.7016, -0.6565,  ..., -0.3714, -0.3864, -0.4164],
         [-0.6565, -0.7016, -0.6865,  ..., -0.3414, -0.3864, -0.4614],
         [-0.6715, -0.7016, -0.6565,  ..., -0.1163, -0.4014, -0.5215]],

        [[-0.5701, -0.5701, -0.5701,  ..., -0.8545, -0.7977, -0.7692],
         [-0.5844, -0.5701, -0.5986,  ..., -0.8688, -0.8261, -0.7692],
         [-0.5133, -0.5701, -0.5133,  ..., -0.7977, -0.8119, -0.8119],
         ...,
         [-0.5701, -0.5417, -0.5417,  ..., -1.4233, -1.4233, -1.4233],
         [-0.6128, -0.6128, -0.4706,  ..., -1.4233, -1.4091, -1.4233],
         [-0.6412, -0.5844, -0.5986,  ..., -1.2954, -1.4660, -1.4091]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat suggestion could you give to someone who wants to make their pizza more interesting and visually appealing? ASSISTANT: To make a pizza more interesting and visually appealing, you can experiment with a variety of colorful and diverse ingredients on top of the pizza. In the image, there are two different types of pizzas on a cutting board. One pizza is covered in colorful veggies, which adds visual interest and increased nutritional value. Incorporating various textures, flavors, and colors using a mix of vegetables, meats, cheeses, and herbs can make a pizza more enticing to the eye and the palate. Additionally, you can try different arrangements of toppings, making patterns or designs, to create a visually engaging presentation. Remember to balance flavors and textures, ensuring no single ingredient overpowers the others, and always use fresh, high-quality ingredients for the best results.</s>"], tensor([], size=(0, 640, 478)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 765), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat suggestion could you give to someone who wants to make their pizza more interesting and visually appealing? ASSISTANT: To make a pizza more interesting and visually appealing, you can experiment with a variety of colorful and diverse ingredients on top of the pizza. In the image, there are two different types of pizzas on a cutting board. One pizza is covered in colorful veggies, which adds visual interest and increased nutritional value. Incorporating various textures, flavors, and colors using a mix of vegetables, meats, cheeses, and herbs can make a pizza more enticing to the eye and the palate. Additionally, you can try different arrangements of toppings, making patterns or designs, to create a visually engaging presentation. Remember to balance flavors and textures, ensuring no single ingredient overpowers the others, and always use fresh, high-quality ingredients for the best results.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat suggestion could you give to someone who wants to make their pizza more interesting and visually appealing? ASSISTANT: To make a pizza more interesting and visually appealing, you can experiment with a variety of colorful and diverse ingredients on top of the pizza. In the image, there are two different types of pizzas on a cutting board. One pizza is covered in colorful veggies, which adds visual interest and increased nutritional value. Incorporating various textures, flavors, and colors using a mix of vegetables, meats, cheeses, and herbs can make a pizza more enticing to the eye and the palate. Additionally, you can try different arrangements of toppings, making patterns or designs, to create a visually engaging presentation. Remember to balance flavors and textures, ensuring no single ingredient overpowers the others, and always use fresh, high-quality ingredients for the best results.</s>"], [{}], False, 'vqa')]
[rank0]: Traceback (most recent call last):
[rank0]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 849, in train
[rank0]:     input_dict = next(train_iter)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 124, in __next__
[rank0]:     return next(self.data)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 157, in <genexpr>
[rank0]:     self.data = (x for x in self.dataloader)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 631, in __next__
[rank0]:     data = self._next_data()
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1346, in _next_data
[rank0]:     return self._process_data(data)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1372, in _process_data
[rank0]:     data.reraise()
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/_utils.py", line 705, in reraise
[rank0]:     raise exception
[rank0]: ValueError: Caught ValueError in DataLoader worker process 0.
[rank0]: Original Traceback (most recent call last):
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/worker.py", line 308, in _worker_loop
[rank0]:     data = fetcher.fetch(index)  # type: ignore[possibly-undefined]
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/fetch.py", line 54, in fetch
[rank0]:     return self.collate_fn(data)
[rank0]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/utils/dataset.py", line 53, in collate_fn
[rank0]:     for (
[rank0]: ValueError: too many values to unpack (expected 12)


[rank0]: During handling of the above exception, another exception occurred:

[rank0]: Traceback (most recent call last):
[rank0]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 1282, in <module>
[rank0]:     main(sys.argv[1:])
[rank0]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 743, in main
[rank0]:     train_iter = train(
[rank0]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 852, in train
[rank0]:     input_dict = next(train_iter)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 124, in __next__
[rank0]:     return next(self.data)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 157, in <genexpr>
[rank0]:     self.data = (x for x in self.dataloader)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 631, in __next__
[rank0]:     data = self._next_data()
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1346, in _next_data
[rank0]:     return self._process_data(data)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1372, in _process_data
[rank0]:     data.reraise()
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/_utils.py", line 705, in reraise
[rank0]:     raise exception
[rank0]: ValueError: Caught ValueError in DataLoader worker process 0.
[rank0]: Original Traceback (most recent call last):
[rank0]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/plum_train_ds.py", line 849, in train
[rank0]:     input_dict = next(train_iter)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 124, in __next__
[rank0]:     return next(self.data)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/deepspeed/runtime/dataloader.py", line 157, in <genexpr>
[rank0]:     self.data = (x for x in self.dataloader)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 631, in __next__
[rank0]:     data = self._next_data()
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1346, in _next_data
[rank0]:     return self._process_data(data)
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/dataloader.py", line 1372, in _process_data
[rank0]:     data.reraise()
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/_utils.py", line 705, in reraise
[rank0]:     raise exception
[rank0]: ValueError: Caught ValueError in DataLoader worker process 0.
[rank0]: Original Traceback (most recent call last):
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/worker.py", line 308, in _worker_loop
[rank0]:     data = fetcher.fetch(index)  # type: ignore[possibly-undefined]
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/fetch.py", line 54, in fetch
[rank0]:     return self.collate_fn(data)
[rank0]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/utils/dataset.py", line 53, in collate_fn
[rank0]:     for (
[rank0]: ValueError: too many values to unpack (expected 12)


[rank0]: During handling of the above exception, another exception occurred:

[rank0]: Traceback (most recent call last):
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/worker.py", line 308, in _worker_loop
[rank0]:     data = fetcher.fetch(index)  # type: ignore[possibly-undefined]
[rank0]:   File "/shared/nas/data/m1/jk100/.conda/envs/llava/lib/python3.10/site-packages/torch/utils/data/_utils/fetch.py", line 54, in fetch
[rank0]:     return self.collate_fn(data)
[rank0]:   File "/shared/nas2/jk100/partonomy_private/src/models/PLUM/utils/dataset.py", line 53, in collate_fn
[rank0]:     for (
[rank0]: ValueError: too many values to unpack (expected 12)

>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2010_004966.jpg', tensor([[[1.0331, 1.0844, 1.1872,  ..., 0.8104, 0.7933, 0.7933],
         [1.0502, 1.0844, 1.1700,  ..., 0.8276, 0.8104, 0.8104],
         [1.1015, 1.1015, 1.1529,  ..., 0.8447, 0.8276, 0.8276],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.3431, 1.3957, 1.5007,  ..., 1.3782, 1.3957, 1.3957],
         [1.3606, 1.3957, 1.4832,  ..., 1.3957, 1.4132, 1.4132],
         [1.4132, 1.4132, 1.4657,  ..., 1.4307, 1.4307, 1.4307],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.0997, 2.1520, 2.2566,  ..., 2.2566, 2.2566, 2.2566],
         [2.1171, 2.1520, 2.2391,  ..., 2.2566, 2.2740, 2.2740],
         [2.1694, 2.1694, 2.2217,  ..., 2.2740, 2.2914, 2.2914],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 0.8792,  0.8792,  0.8938,  ...,  0.8355,  0.8647,  0.8209],
         [ 0.8938,  0.9376,  0.9230,  ...,  0.8501,  0.8501,  0.8647],
         [ 0.9522,  0.9522,  0.8938,  ...,  0.8647,  0.8938,  0.8647],
         ...,
         [-0.9018, -0.8872, -1.2083,  ..., -1.0477, -1.1645, -0.8580],
         [-0.7412, -1.0039, -0.9456,  ..., -0.7120, -0.7266, -0.6682],
         [-1.1207, -1.0623, -1.6901,  ..., -0.6974, -0.6682, -0.4346]],

        [[ 1.1594,  1.1594,  1.1744,  ...,  1.2194,  1.2495,  1.2044],
         [ 1.1744,  1.2194,  1.2044,  ...,  1.2344,  1.2344,  1.2495],
         [ 1.2344,  1.2344,  1.1744,  ...,  1.2495,  1.2795,  1.2495],
         ...,
         [-0.9567, -0.8666, -1.1968,  ..., -1.0617, -1.0767, -0.8366],
         [-0.7916, -1.0017, -0.9117,  ..., -0.7166, -0.6565, -0.6565],
         [-1.2268, -1.1368, -1.7071,  ..., -0.6865, -0.6415, -0.4614]],

        [[ 1.7335,  1.7335,  1.7477,  ...,  1.9042,  1.9326,  1.8899],
         [ 1.7477,  1.7904,  1.7762,  ...,  1.9184,  1.9184,  1.9326],
         [ 1.8046,  1.8046,  1.7477,  ...,  1.9326,  1.9610,  1.9326],
         ...,
         [-1.0536, -0.9967, -1.1958,  ..., -1.0821, -1.2811, -1.2243],
         [-0.8972, -1.1105, -0.9399,  ..., -0.8688, -0.8972, -1.0394],
         [-1.1958, -1.0394, -1.4518,  ..., -0.9541, -0.9825, -0.8972]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is car wheel in this image? Please elaborate your answer and explain why. ASSISTANT: Result: car wheel.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the the wheel of the car in this picture? ASSISTANT: The region corresponding to the the wheel of the car is shown.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the the wheel of the car in this image. ASSISTANT: Below you can see the the wheel of the car.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the the body of the aeroplane from the image below? ASSISTANT: You can see the the body of the aeroplane in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the car license plate in this image? ASSISTANT: The car license plate portion is displayed below.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (733, 1024), ['<image>\nWhat is car wheel in this image? Please elaborate your answer and explain why.', '<image>\nCould you identify and segment out the the wheel of the car in this picture?', '<image>\nPlease segment the the wheel of the car in this image.', '<image>\nWould you please extract the the body of the aeroplane from the image below?', '<image>\nCan you segment the car license plate in this image?'], ['car wheel', 'the wheel of the car', 'the wheel of the car', 'the body of the aeroplane', 'car license plate'], [{'car wheel': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the wheel of the car': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the wheel of the car': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the body of the aeroplane': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'car license plate': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/saiapr_tc-12/08/images/8680.jpg', tensor([[[ 1.3242,  1.3242,  1.3242,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.3242,  1.3242,  1.3242,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.3242,  1.3242,  1.3242,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.3302, -1.3130, -1.2959,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2617, -1.2617, -1.2788,  ...,  0.0000,  0.0000,  0.0000],
         [-1.2445, -1.2445, -1.2617,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.6933,  1.6933,  1.6933,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.6933,  1.6933,  1.6933,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.6933,  1.6933,  1.6933,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.1779, -1.1604, -1.1604,  ...,  0.0000,  0.0000,  0.0000],
         [-1.1078, -1.1078, -1.1254,  ...,  0.0000,  0.0000,  0.0000],
         [-1.0903, -1.0903, -1.1078,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.1868,  2.1868,  2.1868,  ...,  0.0000,  0.0000,  0.0000],
         [ 2.1868,  2.1868,  2.1868,  ...,  0.0000,  0.0000,  0.0000],
         [ 2.1868,  2.1868,  2.1868,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.6999, -1.6824, -1.6476,  ...,  0.0000,  0.0000,  0.0000],
         [-1.6302, -1.6302, -1.6127,  ...,  0.0000,  0.0000,  0.0000],
         [-1.6127, -1.6127, -1.5953,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.3610,  1.3756,  1.4048,  ...,  1.2734,  1.2588,  1.2442],
         [ 1.3902,  1.4194,  1.4632,  ...,  1.2442,  1.2296,  1.2150],
         [ 1.4340,  1.4632,  1.5070,  ...,  1.2150,  1.2004,  1.1858],
         ...,
         [-0.7558, -0.8726, -0.9164,  ..., -0.6974, -0.7704, -0.8142],
         [-0.8872, -0.8726, -0.8434,  ..., -0.7704, -0.9164, -0.9748],
         [-0.8434, -0.7996, -0.7850,  ..., -0.8434, -1.0185, -1.0915]],

        [[ 1.5346,  1.5496,  1.5796,  ...,  1.4596,  1.4446,  1.4295],
         [ 1.5646,  1.5946,  1.6247,  ...,  1.4295,  1.4145,  1.3995],
         [ 1.6096,  1.6397,  1.6847,  ...,  1.3995,  1.3845,  1.3695],
         ...,
         [-0.5665, -0.6865, -0.7316,  ..., -0.3414, -0.4314, -0.4614],
         [-0.7466, -0.7166, -0.6865,  ..., -0.3864, -0.5215, -0.5965],
         [-0.6865, -0.6415, -0.6265,  ..., -0.4614, -0.6415, -0.7316]],

        [[ 1.7620,  1.7762,  1.8046,  ...,  1.6055,  1.5913,  1.5771],
         [ 1.7904,  1.8188,  1.8615,  ...,  1.5771,  1.5629,  1.5487],
         [ 1.8188,  1.8331,  1.8757,  ...,  1.5487,  1.5344,  1.5202],
         ...,
         [-1.3665, -1.4376, -1.4518,  ..., -1.1674, -1.2669, -1.2811],
         [-1.4233, -1.4091, -1.3807,  ..., -1.1389, -1.2954, -1.3380],
         [-1.3807, -1.3380, -1.3238,  ..., -1.1816, -1.3238, -1.3807]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the top of image in this picture? ASSISTANT: This is the selected top of image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is sky in this image? Please elaborate your answer and explain why. ASSISTANT: This view focuses on the sky.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the sky in this image. ASSISTANT: Here is the sky you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the sky in this image? Please elaborate your answer and explain why. ASSISTANT: Below you can see the the sky.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is sky in this image? Please elaborate your answer and explain why. ASSISTANT: sky.</s>"], tensor([[[1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         [1, 1, 1,  ..., 1, 1, 1],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 768), ['<image>\nCould you isolate the top of image in this picture?', '<image>\nWhat is sky in this image? Please elaborate your answer and explain why.', '<image>\nPlease highlight the sky in this image.', '<image>\nWhat is the sky in this image? Please elaborate your answer and explain why.', '<image>\nWhat is sky in this image? Please elaborate your answer and explain why.'], ['top of image', 'sky', 'sky', 'the sky', 'sky'], [{'top of image': tensor([[1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'sky': tensor([[1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'sky': tensor([[1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the sky': tensor([[1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'sky': tensor([[1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        [1, 1, 1,  ..., 1, 1, 1],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000001756.jpg', tensor([[[ 2.0092,  2.0092,  1.9920,  ..., -1.1932, -1.1418, -1.1075],
         [ 2.0263,  2.0263,  2.0092,  ..., -1.2274, -1.1589, -1.1247],
         [ 2.0434,  2.0434,  2.0263,  ..., -1.2617, -1.1932, -1.1418],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.9034,  1.9034,  1.9034,  ..., -1.1604, -1.1779, -1.1779],
         [ 1.9034,  1.9034,  1.9209,  ..., -1.1954, -1.2129, -1.2129],
         [ 1.9209,  1.9209,  1.9384,  ..., -1.2304, -1.2479, -1.2479],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.7163,  1.7163,  1.7163,  ..., -0.9504, -0.9330, -0.9330],
         [ 1.7511,  1.7511,  1.7511,  ..., -0.9853, -0.9678, -0.9504],
         [ 1.7860,  1.7860,  1.7860,  ..., -1.0201, -1.0027, -0.9853],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 9.8144e-01,  1.0398e+00,  1.0982e+00,  ..., -9.7475e-01,
          -9.8935e-01, -9.7475e-01],
         [ 1.0982e+00,  1.1566e+00,  1.1858e+00,  ..., -9.3096e-01,
          -9.4555e-01, -9.4555e-01],
         [ 1.1858e+00,  1.2296e+00,  1.2734e+00,  ..., -9.1636e-01,
          -9.1636e-01, -9.3096e-01],
         ...,
         [ 8.2086e-01,  9.0845e-01,  9.2304e-01,  ...,  1.2004e+00,
           9.3764e-01, -5.0760e-01],
         [ 7.3327e-01,  7.6246e-01,  8.9385e-01,  ..., -8.4247e-02,
           7.3327e-01, -9.3096e-01],
         [ 7.3327e-01,  7.3327e-01,  8.2086e-01,  ...,  2.8071e-01,
           3.8290e-01, -1.3397e+00]],

        [[ 7.8422e-01,  8.1423e-01,  8.5925e-01,  ..., -1.0167e+00,
          -1.0317e+00, -1.0317e+00],
         [ 8.7426e-01,  9.0428e-01,  9.3429e-01,  ..., -1.0017e+00,
          -1.0167e+00, -1.0017e+00],
         [ 9.6431e-01,  1.0093e+00,  1.0694e+00,  ..., -9.5669e-01,
          -9.5669e-01, -9.8670e-01],
         ...,
         [ 9.9432e-01,  1.1144e+00,  1.1444e+00,  ...,  1.1894e+00,
           7.2418e-01, -7.3157e-01],
         [ 8.8927e-01,  9.3429e-01,  1.0994e+00,  ..., -2.0630e-01,
           6.4915e-01, -1.0317e+00],
         [ 9.0428e-01,  9.1929e-01,  9.9432e-01,  ...,  3.1897e-01,
           3.3398e-01, -1.3619e+00]],

        [[ 6.2435e-01,  6.2435e-01,  6.8123e-01,  ..., -8.2610e-01,
          -8.4032e-01, -8.5454e-01],
         [ 7.2389e-01,  7.2389e-01,  7.5233e-01,  ..., -8.4032e-01,
          -8.4032e-01, -8.4032e-01],
         [ 7.9499e-01,  8.0921e-01,  8.8031e-01,  ..., -8.1188e-01,
          -8.4032e-01, -8.6876e-01],
         ...,
         [ 1.3069e+00,  1.3922e+00,  1.3922e+00,  ...,  1.1647e+00,
           1.1221e+00, -1.3329e-03],
         [ 1.2358e+00,  1.2358e+00,  1.3922e+00,  ...,  3.8261e-01,
           9.2297e-01, -5.7014e-01],
         [ 1.1647e+00,  1.1505e+00,  1.2785e+00,  ...,  4.6793e-01,
           5.5325e-01, -9.2564e-01]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nHow many people can be seen in the image? ASSISTANT: There is a large group of people visible in the image, sitting at tables.</s>USER: Are the people sitting with laptops, bags, or both? ASSISTANT: In the image, the people are sitting with both laptops and bags at the tables.</s>USER: Is there a man sitting alone at a table or are all tables occupied by multiple people? ASSISTANT: There is a man sitting alone at a table in front of a crowd of people, separate from the large group at the other tables. He might be a speaker or presenter, given his position relative to the crowd.</s>USER: What might be the purpose of this gathering? ASSISTANT: Based on the image, it appears to be a professional or academic gathering where people are working, collaborating or attending a presentation. The presence of laptops and bags suggests that it could be a conference, workshop, or seminar. The man sitting alone at a table in front of the crowd could be a speaker, presenter, or someone leading the event. The participants' attention may be focused on him or on their laptops for note-taking, research, or other tasks relevant to the occasion.</s>"], tensor([], size=(0, 480, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nHow many people can be seen in the image? ASSISTANT: There is a large group of people visible in the image, sitting at tables.</s>USER: Are the people sitting with laptops, bags, or both? ASSISTANT: In the image, the people are sitting with both laptops and bags at the tables.</s>USER: Is there a man sitting alone at a table or are all tables occupied by multiple people? ASSISTANT: There is a man sitting alone at a table in front of a crowd of people, separate from the large group at the other tables. He might be a speaker or presenter, given his position relative to the crowd.</s>USER: What might be the purpose of this gathering? ASSISTANT: Based on the image, it appears to be a professional or academic gathering where people are working, collaborating or attending a presentation. The presence of laptops and bags suggests that it could be a conference, workshop, or seminar. The man sitting alone at a table in front of the crowd could be a speaker, presenter, or someone leading the event. The participants' attention may be focused on him or on their laptops for note-taking, research, or other tasks relevant to the occasion.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nHow many people can be seen in the image? ASSISTANT: There is a large group of people visible in the image, sitting at tables.</s>USER: Are the people sitting with laptops, bags, or both? ASSISTANT: In the image, the people are sitting with both laptops and bags at the tables.</s>USER: Is there a man sitting alone at a table or are all tables occupied by multiple people? ASSISTANT: There is a man sitting alone at a table in front of a crowd of people, separate from the large group at the other tables. He might be a speaker or presenter, given his position relative to the crowd.</s>USER: What might be the purpose of this gathering? ASSISTANT: Based on the image, it appears to be a professional or academic gathering where people are working, collaborating or attending a presentation. The presence of laptops and bags suggests that it could be a conference, workshop, or seminar. The man sitting alone at a table in front of the crowd could be a speaker, presenter, or someone leading the event. The participants' attention may be focused on him or on their laptops for note-taking, research, or other tasks relevant to the occasion.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000525179.jpg', tensor([[[ 0.1939,  0.1939,  0.2111,  ..., -1.6213, -1.6042, -1.6042],
         [ 0.2111,  0.2111,  0.2282,  ..., -1.6042, -1.6042, -1.6042],
         [ 0.2282,  0.2282,  0.2453,  ..., -1.5870, -1.5870, -1.5870],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.0126,  0.0301,  0.0476,  ..., -1.6856, -1.7031, -1.7031],
         [ 0.0126,  0.0301,  0.0476,  ..., -1.6681, -1.6856, -1.7031],
         [ 0.0301,  0.0301,  0.0476,  ..., -1.6331, -1.6681, -1.7031],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.0605,  0.0953,  0.1302,  ..., -1.4036, -1.4036, -1.4036],
         [ 0.0953,  0.1128,  0.1302,  ..., -1.4210, -1.4384, -1.4384],
         [ 0.1302,  0.1302,  0.1476,  ..., -1.4559, -1.4733, -1.4733],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.2223,  0.2515,  0.2223,  ..., -0.8142, -0.8434, -0.8580],
         [ 0.2515,  0.2807,  0.2807,  ..., -0.7996, -0.8434, -0.8434],
         [ 0.2661,  0.2515,  0.2369,  ..., -0.7996, -0.7996, -0.7850],
         ...,
         [ 0.3099, -0.3470, -1.2667,  ..., -1.7485, -1.7485, -1.7485],
         [ 0.5289, -0.5806, -1.3105,  ..., -1.7485, -1.7485, -1.7485],
         [ 0.4559, -0.7850, -1.3835,  ..., -1.7339, -1.7339, -1.7485]],

        [[ 0.0638,  0.1089,  0.0789,  ..., -0.8967, -0.8967, -0.9417],
         [ 0.0939,  0.1239,  0.1239,  ..., -0.8816, -0.8967, -0.9267],
         [ 0.1089,  0.0939,  0.0789,  ..., -0.8967, -0.8816, -0.9267],
         ...,
         [ 0.3190, -0.2213, -1.2118,  ..., -1.7071, -1.7071, -1.7071],
         [ 0.5591, -0.3864, -1.2568,  ..., -1.7071, -1.7071, -1.7071],
         [ 0.4991, -0.5365, -1.2869,  ..., -1.7221, -1.7221, -1.7071]],

        [[ 0.1266,  0.1551,  0.1409,  ..., -0.7123, -0.7123, -0.7266],
         [ 0.1551,  0.1977,  0.1835,  ..., -0.7266, -0.7123, -0.6981],
         [ 0.1693,  0.1266,  0.1124,  ..., -0.7123, -0.6981, -0.6981],
         ...,
         [ 0.4395,  0.0129, -0.9256,  ..., -1.4660, -1.4660, -1.4660],
         [ 0.6244, -0.2004, -0.9825,  ..., -1.4518, -1.4518, -1.4802],
         [ 0.5959, -0.3711, -1.0394,  ..., -1.4376, -1.4518, -1.4802]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the man in this picture? ASSISTANT: You can see the man in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the girl in this picture? ASSISTANT: This view focuses on the girl.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the black suit from the provided image. ASSISTANT: Below you can see the black suit.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the woman on right in this image. ASSISTANT: This is the selected woman on right.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the owoman in this picture. ASSISTANT: Here's where the owoman appears in the image.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nCould you identify and segment out the man in this picture?', '<image>\nCould you isolate the girl in this picture?', '<image>\nSegment the black suit from the provided image.', '<image>\nPlease highlight the woman on right in this image.', '<image>\nPlease point out the owoman in this picture.'], ['man', 'girl', 'black suit', 'woman on right', 'owoman'], [{'man': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'girl': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'black suit': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'woman on right': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'owoman': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000217302.jpg', tensor([[[ 2.1119,  2.1119,  2.1290,  ..., -0.1143, -0.1314, -0.1486],
         [ 2.1119,  2.1119,  2.1290,  ..., -0.0972, -0.1143, -0.1143],
         [ 2.1119,  2.1119,  2.1290,  ..., -0.0801, -0.0801, -0.0801],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.3585,  2.3585,  2.3761,  ...,  0.2402,  0.2227,  0.2227],
         [ 2.3585,  2.3585,  2.3761,  ...,  0.2577,  0.2402,  0.2402],
         [ 2.3585,  2.3585,  2.3761,  ...,  0.2752,  0.2752,  0.2752],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.6226,  2.6226,  2.6400,  ...,  0.1476,  0.1476,  0.1476],
         [ 2.6226,  2.6226,  2.6400,  ...,  0.1651,  0.1651,  0.1651],
         [ 2.6226,  2.6226,  2.6400,  ...,  0.1825,  0.1999,  0.1999],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.8135,  1.8135,  1.8281,  ...,  0.0471,  0.0617,  0.0325],
         [ 1.8281,  1.8281,  1.8281,  ...,  0.0471,  0.0617,  0.0179],
         [ 1.8135,  1.8135,  1.8281,  ...,  0.0471,  0.0325,  0.0617],
         ...,
         [-0.3470, -0.6244, -0.3324,  ...,  0.2807,  0.3975,  0.3537],
         [-0.1426, -0.1426,  0.6019,  ...,  0.4121,  0.4559,  0.4559],
         [ 0.9814,  1.1712,  1.1566,  ...,  0.3683,  0.3829,  0.4413]],

        [[ 1.9998,  2.0149,  2.0299,  ...,  0.3490,  0.3490,  0.2890],
         [ 2.0149,  2.0299,  2.0299,  ...,  0.3340,  0.3490,  0.3040],
         [ 1.9998,  2.0149,  2.0299,  ...,  0.3340,  0.3040,  0.3040],
         ...,
         [-1.3619, -1.3169, -0.6565,  ...,  0.4090,  0.5591,  0.5891],
         [-1.0317, -0.5365,  0.5441,  ...,  0.5891,  0.6341,  0.6942],
         [ 0.9793,  1.1594,  1.1444,  ...,  0.5741,  0.5891,  0.6792]],

        [[ 2.1459,  2.1459,  2.1459,  ...,  0.2404,  0.2404,  0.2262],
         [ 2.1459,  2.1459,  2.1459,  ...,  0.2546,  0.2546,  0.2262],
         [ 2.1459,  2.1459,  2.1459,  ...,  0.2404,  0.2546,  0.2688],
         ...,
         [-1.3380, -0.9967, -0.5133,  ...,  0.7523,  0.8519,  0.8234],
         [-0.8688, -0.2857,  0.5390,  ...,  0.8377,  0.9372,  0.8945],
         [ 1.0083,  1.1647,  1.1363,  ...,  0.8234,  0.8519,  0.9088]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat potential safety concerns could be present for the bus and other road users in this situation? ASSISTANT: During the rain and on wet streets, safety concerns for the bus and other road users include the increased risk of accidents due to slippery roads, reduced visibility, and longer braking distances. In the image, the red city bus is driving on a rainy day, and the streets are wet. Drivers, including the bus driver, need to be cautious and maintain a safe distance from each other, adjust their speed according to the road conditions, and ensure their headlights and windshield wipers are functioning properly. Pedestrians and cyclists should also exercise additional care while navigating through the wet streets. It is crucial for all road users to remain vigilant and take appropriate safety measures to avoid potential accidents and ensure a secure journey for everyone during such weather conditions.</s>"], tensor([], size=(0, 480, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat potential safety concerns could be present for the bus and other road users in this situation? ASSISTANT: During the rain and on wet streets, safety concerns for the bus and other road users include the increased risk of accidents due to slippery roads, reduced visibility, and longer braking distances. In the image, the red city bus is driving on a rainy day, and the streets are wet. Drivers, including the bus driver, need to be cautious and maintain a safe distance from each other, adjust their speed according to the road conditions, and ensure their headlights and windshield wipers are functioning properly. Pedestrians and cyclists should also exercise additional care while navigating through the wet streets. It is crucial for all road users to remain vigilant and take appropriate safety measures to avoid potential accidents and ensure a secure journey for everyone during such weather conditions.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat potential safety concerns could be present for the bus and other road users in this situation? ASSISTANT: During the rain and on wet streets, safety concerns for the bus and other road users include the increased risk of accidents due to slippery roads, reduced visibility, and longer braking distances. In the image, the red city bus is driving on a rainy day, and the streets are wet. Drivers, including the bus driver, need to be cautious and maintain a safe distance from each other, adjust their speed according to the road conditions, and ensure their headlights and windshield wipers are functioning properly. Pedestrians and cyclists should also exercise additional care while navigating through the wet streets. It is crucial for all road users to remain vigilant and take appropriate safety measures to avoid potential accidents and ensure a secure journey for everyone during such weather conditions.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000455865.jpg', tensor([[[-1.8268, -1.8610, -1.9124,  ...,  0.6563,  0.6392,  0.6392],
         [-1.8439, -1.8610, -1.8782,  ...,  0.6563,  0.6392,  0.6392],
         [-1.8782, -1.8610, -1.8439,  ...,  0.6563,  0.6563,  0.6563],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.8957, -1.9132, -1.9307,  ..., -0.5301, -0.5301, -0.5301],
         [-1.8957, -1.8957, -1.8957,  ..., -0.5301, -0.5301, -0.5301],
         [-1.8957, -1.8782, -1.8606,  ..., -0.5476, -0.5301, -0.5301],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.6127, -1.6302, -1.6650,  ..., -1.6302, -1.6824, -1.7173],
         [-1.6127, -1.6127, -1.6127,  ..., -1.6127, -1.6476, -1.6824],
         [-1.5953, -1.5779, -1.5604,  ..., -1.5779, -1.6127, -1.6302],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.6609, -1.6609, -1.6755,  ...,  0.2807,  0.2953,  0.3245],
         [-1.6317, -1.6609, -1.6755,  ...,  0.4705,  0.4559,  0.4997],
         [-1.6755, -1.6901, -1.7047,  ...,  0.5581,  0.5727,  0.5727],
         ...,
         [ 0.9084,  0.9814,  0.9668,  ...,  0.5143,  0.5873,  0.6457],
         [ 0.9084,  0.9522,  0.9668,  ...,  0.4705,  0.5435,  0.6019],
         [ 0.9230,  0.9230,  0.9522,  ...,  0.4267,  0.5289,  0.5581]],

        [[-1.6771, -1.6621, -1.6771,  ..., -0.7166, -0.6715, -0.6715],
         [-1.6621, -1.6771, -1.6771,  ..., -0.5965, -0.6115, -0.5815],
         [-1.6470, -1.6771, -1.6621,  ..., -0.5215, -0.5065, -0.5065],
         ...,
         [ 0.0188,  0.1089,  0.0939,  ..., -0.5815, -0.5065, -0.4014],
         [ 0.0188,  0.0939,  0.0939,  ..., -0.6115, -0.5365, -0.4614],
         [ 0.0338,  0.0638,  0.0939,  ..., -0.6565, -0.5815, -0.5215]],

        [[-1.3522, -1.3522, -1.3665,  ..., -1.3665, -1.3238, -1.3949],
         [-1.3380, -1.3522, -1.3665,  ..., -1.3665, -1.3522, -1.3380],
         [-1.3522, -1.3665, -1.3665,  ..., -1.2954, -1.3522, -1.3380],
         ...,
         [-1.0821, -1.1247, -1.1247,  ..., -1.3096, -1.2811, -1.3522],
         [-1.0821, -1.1247, -1.1247,  ..., -1.3522, -1.2811, -1.3238],
         [-1.0252, -1.0963, -1.0821,  ..., -1.4091, -1.3238, -1.3949]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat considerations should someone have while eating this pizza? ASSISTANT: While eating this pizza, which appears to be a Fettuccini Alfredo pizza or a white pizza with chicken, one should consider a few things. The pizza may be quite rich and greasy due to the presence of sausage, chicken, and possibly creamy ingredients like Alfredo sauce. To overcome this, they may choose to blot excess grease using a paper napkin, or be mindful of their portion size when consuming the pizza. Also, anyone with dietary restrictions, such as lactose intolerance, should be cautious when eating this type of pizza due to potential creamy sauce ingredients. Finally, it is essential to ensure proper hygiene and etiquette, especially if sharing the pizza with others at the dining table. Using clean hands or utensils and taking turns to serve the slices can make the experience more enjoyable for everyone.</s>"], tensor([], size=(0, 428, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (685, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat considerations should someone have while eating this pizza? ASSISTANT: While eating this pizza, which appears to be a Fettuccini Alfredo pizza or a white pizza with chicken, one should consider a few things. The pizza may be quite rich and greasy due to the presence of sausage, chicken, and possibly creamy ingredients like Alfredo sauce. To overcome this, they may choose to blot excess grease using a paper napkin, or be mindful of their portion size when consuming the pizza. Also, anyone with dietary restrictions, such as lactose intolerance, should be cautious when eating this type of pizza due to potential creamy sauce ingredients. Finally, it is essential to ensure proper hygiene and etiquette, especially if sharing the pizza with others at the dining table. Using clean hands or utensils and taking turns to serve the slices can make the experience more enjoyable for everyone.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat considerations should someone have while eating this pizza? ASSISTANT: While eating this pizza, which appears to be a Fettuccini Alfredo pizza or a white pizza with chicken, one should consider a few things. The pizza may be quite rich and greasy due to the presence of sausage, chicken, and possibly creamy ingredients like Alfredo sauce. To overcome this, they may choose to blot excess grease using a paper napkin, or be mindful of their portion size when consuming the pizza. Also, anyone with dietary restrictions, such as lactose intolerance, should be cautious when eating this type of pizza due to potential creamy sauce ingredients. Finally, it is essential to ensure proper hygiene and etiquette, especially if sharing the pizza with others at the dining table. Using clean hands or utensils and taking turns to serve the slices can make the experience more enjoyable for everyone.</s>"], [{}], False, 'vqa')]
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000291854.jpg', tensor([[[ 2.2489,  2.2489,  2.2489,  ..., -1.4329, -0.9877, -0.6109],
         [ 2.2489,  2.2489,  2.2489,  ..., -1.6555, -1.3473, -1.0904],
         [ 2.2489,  2.2489,  2.2489,  ..., -1.9295, -1.7925, -1.6898],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.4286,  2.4286,  2.4286,  ..., -1.3704, -0.8803, -0.4776],
         [ 2.4286,  2.4286,  2.4286,  ..., -1.5980, -1.2479, -0.9678],
         [ 2.4286,  2.4286,  2.4286,  ..., -1.8782, -1.7031, -1.5805],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.6400,  2.6400,  2.6400,  ..., -1.0550, -0.5670, -0.1835],
         [ 2.6400,  2.6400,  2.6400,  ..., -1.2816, -0.9330, -0.6715],
         [ 2.6400,  2.6400,  2.6400,  ..., -1.5604, -1.4036, -1.2816],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 1.7698,  1.9303,  1.4194,  ..., -1.3981, -1.0769, -1.0623],
         [ 1.9157,  1.7552,  1.1274,  ..., -1.3397, -1.1061, -1.1207],
         [ 1.2734,  1.4194,  0.3829,  ..., -1.1791, -1.0331, -1.2521],
         ...,
         [ 0.2223,  0.0033,  0.1055,  ..., -0.0259, -1.0039, -1.2375],
         [-0.0550, -0.0988,  0.1785,  ..., -0.1718, -0.1426, -0.8142],
         [-0.2448, -0.3178,  0.1055,  ..., -0.1426,  0.0033, -0.0842]],

        [[ 2.0149,  2.0449,  1.6247,  ..., -1.1668, -0.7916, -0.7466],
         [ 2.0149,  1.9398,  1.5196,  ..., -1.0767, -0.8216, -0.8516],
         [ 1.6847,  1.6247,  0.8893,  ..., -0.8816, -0.7316, -1.0017],
         ...,
         [ 0.2289, -0.1463,  0.0789,  ..., -0.0712, -0.9417, -1.0017],
         [-0.2813, -0.4764,  0.0789,  ..., -0.4914, -0.6565, -0.8666],
         [-0.5815, -0.5965, -0.1163,  ..., -0.6715, -0.5665, -0.5965]],

        [[ 2.0890,  2.0890,  1.8046,  ..., -1.1532, -0.9256, -1.1816],
         [ 2.0890,  2.0748,  1.4491,  ..., -1.1105, -1.1247, -1.0821],
         [ 1.6909,  1.8615,  0.7523,  ..., -1.0963, -1.1674, -1.1247],
         ...,
         [ 0.2688, -0.0582,  0.1266,  ..., -0.0724, -0.8403, -0.9541],
         [-0.2715, -0.4137,  0.0982,  ..., -0.3711, -0.4564, -0.7977],
         [-0.5986, -0.6412, -0.0582,  ..., -0.5275, -0.4848, -0.4990]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the giraffe in the image doing? ASSISTANT: The giraffe in the image is stretching its neck and reaching its head over a wooden fence to interact with a group of people standing on the other side of the fence.</s>USER: Where is the giraffe most likely located? ASSISTANT: The giraffe is most likely located in a zoo, as it is standing within a fenced area and interacting with onlookers.</s>USER: What material is the fence made of? ASSISTANT: The fence is made of both wood and metal. The giraffe is specifically reaching over the wooden portion of the fence.</s>USER: Are there any other visible animals in the image? ASSISTANT: No, the image only features a single giraffe and a group of people who have gathered to observe and interact with it.</s>USER: What can be said about the relationship between humans and wildlife, based on the image? ASSISTANT: The image of the giraffe stretching its neck over the wooden fence to reach the people on the other side suggests a level of curiosity and interaction between humans and wildlife. It demonstrates how people are fascinated by animals and can have opportunities for safe and respectful interaction in controlled environments like zoos. \n\nZoos serve as a space where people can connect with various animal species and learn about their biology, behavior, and the challenges they face in the wild, such as habitat loss and poaching. Through these interactions, individuals may develop an increased sense of responsibility towards wildlife conservation and might be inspired to support efforts to protect and preserve endangered species and ecosystems.\n\nHowever, it is essential to remember that the well-being of the animals should be a top priority. Zoos and managed wildlife facilities should strive to provide comfortable and enriching living conditions that cater to the natural behaviors and needs of their animal residents. Education and awareness programs at zoos should also emphasize how people can support conservation efforts and make informed choices in their daily lives that could benefit wildlife and habitats around the world.\n\nIn conclusion, the image depicting a giraffe interacting with onlookers over a wooden fence showcases the relationship between humans and wildlife, demonstrating people's fascination and ability to connect with animals in a safe and respectful manner. This interaction can inspire support for wildlife conservation and education, contributing to a better understanding of the natural world and efforts to protect it.</s>"], tensor([], size=(0, 426, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (682, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the giraffe in the image doing? ASSISTANT: The giraffe in the image is stretching its neck and reaching its head over a wooden fence to interact with a group of people standing on the other side of the fence.</s>USER: Where is the giraffe most likely located? ASSISTANT: The giraffe is most likely located in a zoo, as it is standing within a fenced area and interacting with onlookers.</s>USER: What material is the fence made of? ASSISTANT: The fence is made of both wood and metal. The giraffe is specifically reaching over the wooden portion of the fence.</s>USER: Are there any other visible animals in the image? ASSISTANT: No, the image only features a single giraffe and a group of people who have gathered to observe and interact with it.</s>USER: What can be said about the relationship between humans and wildlife, based on the image? ASSISTANT: The image of the giraffe stretching its neck over the wooden fence to reach the people on the other side suggests a level of curiosity and interaction between humans and wildlife. It demonstrates how people are fascinated by animals and can have opportunities for safe and respectful interaction in controlled environments like zoos. \n\nZoos serve as a space where people can connect with various animal species and learn about their biology, behavior, and the challenges they face in the wild, such as habitat loss and poaching. Through these interactions, individuals may develop an increased sense of responsibility towards wildlife conservation and might be inspired to support efforts to protect and preserve endangered species and ecosystems.\n\nHowever, it is essential to remember that the well-being of the animals should be a top priority. Zoos and managed wildlife facilities should strive to provide comfortable and enriching living conditions that cater to the natural behaviors and needs of their animal residents. Education and awareness programs at zoos should also emphasize how people can support conservation efforts and make informed choices in their daily lives that could benefit wildlife and habitats around the world.\n\nIn conclusion, the image depicting a giraffe interacting with onlookers over a wooden fence showcases the relationship between humans and wildlife, demonstrating people's fascination and ability to connect with animals in a safe and respectful manner. This interaction can inspire support for wildlife conservation and education, contributing to a better understanding of the natural world and efforts to protect it.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the giraffe in the image doing? ASSISTANT: The giraffe in the image is stretching its neck and reaching its head over a wooden fence to interact with a group of people standing on the other side of the fence.</s>USER: Where is the giraffe most likely located? ASSISTANT: The giraffe is most likely located in a zoo, as it is standing within a fenced area and interacting with onlookers.</s>USER: What material is the fence made of? ASSISTANT: The fence is made of both wood and metal. The giraffe is specifically reaching over the wooden portion of the fence.</s>USER: Are there any other visible animals in the image? ASSISTANT: No, the image only features a single giraffe and a group of people who have gathered to observe and interact with it.</s>USER: What can be said about the relationship between humans and wildlife, based on the image? ASSISTANT: The image of the giraffe stretching its neck over the wooden fence to reach the people on the other side suggests a level of curiosity and interaction between humans and wildlife. It demonstrates how people are fascinated by animals and can have opportunities for safe and respectful interaction in controlled environments like zoos. \n\nZoos serve as a space where people can connect with various animal species and learn about their biology, behavior, and the challenges they face in the wild, such as habitat loss and poaching. Through these interactions, individuals may develop an increased sense of responsibility towards wildlife conservation and might be inspired to support efforts to protect and preserve endangered species and ecosystems.\n\nHowever, it is essential to remember that the well-being of the animals should be a top priority. Zoos and managed wildlife facilities should strive to provide comfortable and enriching living conditions that cater to the natural behaviors and needs of their animal residents. Education and awareness programs at zoos should also emphasize how people can support conservation efforts and make informed choices in their daily lives that could benefit wildlife and habitats around the world.\n\nIn conclusion, the image depicting a giraffe interacting with onlookers over a wooden fence showcases the relationship between humans and wildlife, demonstrating people's fascination and ability to connect with animals in a safe and respectful manner. This interaction can inspire support for wildlife conservation and education, contributing to a better understanding of the natural world and efforts to protect it.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000091536.jpg', tensor([[[-0.1143, -0.1143, -0.1143,  ...,  0.2111,  0.2282,  0.2282],
         [-0.1143, -0.1143, -0.1143,  ...,  0.2111,  0.2282,  0.2282],
         [-0.1143, -0.1143, -0.1143,  ...,  0.2111,  0.2111,  0.2111],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.6954,  0.6954,  0.6954,  ...,  1.0280,  1.0455,  1.0455],
         [ 0.6954,  0.6954,  0.6954,  ...,  1.0280,  1.0455,  1.0455],
         [ 0.6954,  0.6954,  0.6954,  ...,  1.0280,  1.0280,  1.0280],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.0017,  1.0191,  1.0365,  ...,  1.1934,  1.2282,  1.2457],
         [ 1.0191,  1.0191,  1.0365,  ...,  1.1934,  1.2108,  1.2282],
         [ 1.0365,  1.0365,  1.0365,  ...,  1.1934,  1.1934,  1.1934],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.0325,  0.0179,  0.0179,  ...,  0.2223,  0.2223,  0.2077],
         [-0.0113,  0.0033, -0.0113,  ...,  0.2223,  0.2369,  0.2223],
         [ 0.0033,  0.0033,  0.0033,  ...,  0.2223,  0.2515,  0.2515],
         ...,
         [-0.1134, -0.1280, -0.1864,  ..., -0.2302, -0.2156, -0.1718],
         [-0.0988, -0.0696, -0.0988,  ...,  0.0617, -0.2302, -0.2448],
         [-0.0405, -0.1718, -0.0988,  ..., -0.1426, -0.2156, -0.2010]],

        [[ 0.7392,  0.7242,  0.7242,  ...,  0.9193,  0.9193,  0.9193],
         [ 0.7392,  0.7092,  0.6942,  ...,  0.9343,  0.9493,  0.9343],
         [ 0.7242,  0.7092,  0.7092,  ...,  0.9343,  0.9493,  0.9493],
         ...,
         [-0.2213, -0.2213, -0.2963,  ..., -0.3564, -0.3414, -0.2963],
         [-0.2063, -0.1313, -0.1763,  ..., -0.0262, -0.3714, -0.4164],
         [-0.1763, -0.2963, -0.2363,  ..., -0.2513, -0.3264, -0.3414]],

        [[ 0.9656,  0.9514,  0.9372,  ...,  1.0367,  1.0367,  1.0367],
         [ 0.9372,  0.9230,  0.9230,  ...,  1.0083,  1.0225,  1.0225],
         [ 0.9230,  0.9230,  0.9372,  ...,  1.0367,  1.0510,  1.0510],
         ...,
         [-0.0156,  0.0129, -0.0440,  ..., -0.1151, -0.0440, -0.0440],
         [-0.0013, -0.0013, -0.0156,  ...,  0.0840, -0.1009, -0.1009],
         [ 0.0129, -0.1009, -0.0298,  ..., -0.0156, -0.0724, -0.0582]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the pavement in this picture? ASSISTANT: Displayed here is the pavement.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is sky in this image? Please give some explanation. ASSISTANT: Sure — this shows the sky.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the hill in this image? ASSISTANT: Take a look at the hill here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the bus in this image? ASSISTANT: Here's where the bus appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the house in this picture? ASSISTANT: Here's where the house appears in the image.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[156, 156, 156,  ..., 156, 156, 156],
        [156, 156, 156,  ..., 156, 156, 156],
        [156, 156, 156,  ..., 156, 156, 156],
        ...,
        [148, 148, 148,  ..., 148, 148, 148],
        [148, 148, 148,  ..., 148, 148, 148],
        [148, 148, 148,  ..., 148, 148, 148]]), (768, 1024), ['<image>\nCould you isolate the pavement in this picture?', '<image>\nWhat is sky in this image? Please give some explanation.', '<image>\nCan you segment the hill in this image?', '<image>\nCan you segment the bus in this image?', '<image>\nCould you isolate the house in this picture?'], ['pavement', 'sky', 'hill', 'bus', 'house'], [{'pavement': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'sky': tensor([[ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'hill': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'bus': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'house': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2008_000415.jpg', tensor([[[1.2899, 1.2899, 1.3070,  ..., 0.4508, 0.3652, 0.3309],
         [1.2899, 1.2899, 1.3070,  ..., 0.4508, 0.3823, 0.3481],
         [1.3070, 1.3070, 1.3242,  ..., 0.4679, 0.4337, 0.3994],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.7983, 1.7983, 1.8158,  ..., 0.5903, 0.5028, 0.4678],
         [1.7983, 1.7983, 1.8158,  ..., 0.5903, 0.5203, 0.4853],
         [1.8158, 1.8158, 1.8158,  ..., 0.5903, 0.5378, 0.5203],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.4831, 2.4831, 2.5006,  ..., 0.7925, 0.7228, 0.6879],
         [2.4831, 2.4831, 2.5006,  ..., 0.7925, 0.7402, 0.7228],
         [2.5006, 2.5006, 2.5180,  ..., 0.8099, 0.7925, 0.7751],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.3318,  1.3610,  1.3464,  ...,  0.9522,  0.9376,  0.9522],
         [ 1.3318,  1.3464,  1.3464,  ...,  0.9814,  0.9814,  0.9814],
         [ 1.3464,  1.3464,  1.3610,  ...,  0.9960,  0.9960,  0.9814],
         ...,
         [-0.0696,  0.0179,  0.0471,  ..., -1.2083, -1.2083, -1.1791],
         [ 0.0033, -0.0113,  0.0325,  ..., -1.0477, -1.0331, -1.0185],
         [-0.0259,  0.0033,  0.0033,  ..., -0.9748, -0.9893, -0.9893]],

        [[ 1.6997,  1.6997,  1.6847,  ...,  1.4295,  1.4295,  1.4446],
         [ 1.6997,  1.6847,  1.6847,  ...,  1.4295,  1.4295,  1.4295],
         [ 1.7147,  1.6997,  1.6997,  ...,  1.4145,  1.4295,  1.4145],
         ...,
         [-0.1613, -0.1163, -0.0862,  ..., -1.4820, -1.4669, -1.4369],
         [-0.1763, -0.1913, -0.1012,  ..., -1.3169, -1.3319, -1.3769],
         [-0.2213, -0.1613, -0.1163,  ..., -1.3619, -1.3619, -1.3619]],

        [[ 2.1459,  2.1459,  2.1459,  ...,  1.8757,  1.8473,  1.8615],
         [ 2.1459,  2.1459,  2.1459,  ...,  1.8757,  1.8757,  1.8615],
         [ 2.1459,  2.1459,  2.1459,  ...,  1.8757,  1.8757,  1.8615],
         ...,
         [-0.1151, -0.1293, -0.0724,  ..., -1.2100, -1.1958, -1.1816],
         [-0.1435, -0.0867, -0.0156,  ..., -1.1532, -1.1389, -1.1532],
         [-0.1435, -0.0156, -0.0156,  ..., -1.1247, -1.1247, -1.1247]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the the neck of the person in this image. ASSISTANT: Here's the area for the neck of the person.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the the torso of the person in this picture? ASSISTANT: Result: the torso of the person.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is person nose in this image? Please give some explanation. ASSISTANT: Displayed here is the person nose.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the person head in this image. ASSISTANT: Here's where the person head appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the person arm from the image below? ASSISTANT: Displayed here is the person arm.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [1, 1, 1,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 1,  ..., 0, 0, 0],
         [0, 1, 1,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nPlease highlight the the neck of the person in this image.', '<image>\nCould you isolate the the torso of the person in this picture?', '<image>\nWhat is person nose in this image? Please give some explanation.', '<image>\nPlease highlight the person head in this image.', '<image>\nWould you please extract the person arm from the image below?'], ['the neck of the person', 'the torso of the person', 'person nose', 'person head', 'person arm'], [{'the neck of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the torso of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [1, 1, 1,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person nose': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person head': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person arm': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 1,  ..., 0, 0, 0],
        [0, 1, 1,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/ade20k/images/training/ADE_train_00018042.jpg', tensor([[[ 2.1633,  2.1462,  2.1462,  ...,  0.1768,  0.2453,  0.3138],
         [ 2.0948,  2.0777,  2.0605,  ...,  0.1426,  0.1597,  0.1939],
         [ 2.0434,  2.0092,  1.9920,  ...,  0.1083,  0.0741,  0.0398],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.4111,  2.3936,  2.3936,  ...,  0.6779,  0.7654,  0.8354],
         [ 2.3585,  2.3410,  2.3235,  ...,  0.6429,  0.6779,  0.7129],
         [ 2.3060,  2.2885,  2.2710,  ...,  0.6078,  0.5903,  0.5728],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 2.6400,  2.6226,  2.6051,  ..., -0.0790, -0.0615, -0.0267],
         [ 2.6051,  2.5703,  2.5354,  ..., -0.1312, -0.1487, -0.1487],
         [ 2.5703,  2.5180,  2.5006,  ..., -0.1835, -0.2532, -0.3055],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[1.9011, 1.9157, 1.9157,  ..., 1.9011, 1.9303, 1.9303],
         [1.9157, 1.9303, 1.9303,  ..., 1.9303, 1.9303, 1.9303],
         [1.9303, 1.9303, 1.9303,  ..., 1.9303, 1.9303, 1.9303],
         ...,
         [1.7698, 1.7844, 1.7844,  ..., 1.4778, 1.4778, 1.4486],
         [1.7406, 1.7698, 1.7698,  ..., 1.4048, 1.4048, 1.4194],
         [1.7260, 1.7260, 1.7260,  ..., 1.4486, 1.4632, 1.4194]],

        [[2.0449, 2.0599, 2.0599,  ..., 2.0749, 2.0749, 2.0749],
         [2.0599, 2.0749, 2.0749,  ..., 2.0749, 2.0749, 2.0749],
         [2.0749, 2.0749, 2.0749,  ..., 2.0599, 2.0749, 2.0749],
         ...,
         [1.9098, 1.9098, 1.9098,  ..., 1.6096, 1.6096, 1.5796],
         [1.8648, 1.8948, 1.8948,  ..., 1.5346, 1.5346, 1.5496],
         [1.8498, 1.8498, 1.8498,  ..., 1.5796, 1.5946, 1.5496]],

        [[2.1175, 2.1317, 2.1317,  ..., 2.1317, 2.1459, 2.1459],
         [2.1317, 2.1459, 2.1459,  ..., 2.1317, 2.1459, 2.1459],
         [2.1459, 2.1459, 2.1459,  ..., 2.1459, 2.1459, 2.1459],
         ...,
         [2.0321, 2.0464, 2.0464,  ..., 1.7335, 1.7335, 1.7051],
         [2.0179, 2.0464, 2.0464,  ..., 1.6624, 1.6624, 1.6766],
         [2.0179, 2.0179, 2.0179,  ..., 1.7051, 1.7193, 1.6766]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the sky in this picture. ASSISTANT: You can see the sky in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the person in this image. ASSISTANT: You can see the person in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the road in this image. ASSISTANT: road.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the grass in this image. ASSISTANT: The grass portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is sidewalk in this image? Please elaborate your answer and explain why. ASSISTANT: sidewalk.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False,  True,  True,  ..., False, False, False],
         [False,  True,  True,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False,  True,  True,  ...,  True,  True,  True],
         [False,  True,  True,  ...,  True,  True,  True],
         [False,  True,  True,  ...,  True,  True,  True]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[255, 255, 255,  ..., 255, 255, 255],
        [255,   2,   2,  ...,   4,   4,   4],
        [255,   2,   2,  ...,   4,   4,   4],
        ...,
        [255,   6,   6,  ...,   6,   6,   6],
        [255,   6,   6,  ...,   6,   6,   6],
        [255,   6,   6,  ...,   6,   6,   6]]), (768, 1024), ['<image>\nPlease point out the sky in this picture.', '<image>\nPlease highlight the person in this image.', '<image>\nPlease segment the road in this image.', '<image>\nPlease highlight the grass in this image.', '<image>\nWhat is sidewalk in this image? Please elaborate your answer and explain why.'], ['sky', 'person', 'road', 'grass', 'sidewalk'], [{'sky': tensor([[False, False, False,  ..., False, False, False],
        [False,  True,  True,  ..., False, False, False],
        [False,  True,  True,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'person': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'road': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False,  True,  True,  ...,  True,  True,  True],
        [False,  True,  True,  ...,  True,  True,  True],
        [False,  True,  True,  ...,  True,  True,  True]])}, {'grass': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'sidewalk': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000212610.jpg', tensor([[[-0.7137, -0.6794, -0.6281,  ...,  0.8618,  0.8789,  0.8789],
         [-0.6794, -0.6623, -0.6281,  ...,  0.8618,  0.8618,  0.8618],
         [-0.6452, -0.6281, -0.6281,  ...,  0.8618,  0.8447,  0.8447],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.9153, -0.8627, -0.7927,  ...,  0.9405,  0.9580,  0.9755],
         [-0.8803, -0.8452, -0.7927,  ...,  0.9405,  0.9580,  0.9580],
         [-0.8277, -0.8102, -0.8102,  ...,  0.9580,  0.9405,  0.9405],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.8981, -0.8807, -0.8458,  ...,  0.6531,  0.6531,  0.6531],
         [-0.8981, -0.8633, -0.7936,  ...,  0.6356,  0.6356,  0.6356],
         [-0.9156, -0.8284, -0.7238,  ...,  0.6008,  0.6182,  0.6182],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.2077,  0.2369,  0.2515,  ...,  1.0106,  0.4997, -0.3324],
         [ 0.2369,  0.2369,  0.2515,  ...,  1.1712,  1.1858,  1.0252],
         [ 0.2077,  0.2223,  0.2515,  ...,  1.2004,  1.1858,  1.1420],
         ...,
         [ 0.6165,  0.5727,  0.4851,  ...,  0.2515,  0.5143,  0.5873],
         [ 0.6019,  0.6457,  0.6311,  ...,  0.3099,  0.5581,  0.6311],
         [ 0.5581,  0.4705,  0.3537,  ...,  0.3391,  0.4997,  0.5143]],

        [[ 0.0789,  0.0939,  0.0939,  ...,  0.5141,  0.0038, -0.6415],
         [ 0.1389,  0.1089,  0.0939,  ...,  0.7692,  0.7092,  0.4991],
         [ 0.1089,  0.1539,  0.1089,  ...,  0.7092,  0.6792,  0.6942],
         ...,
         [-0.1763, -0.1763, -0.1763,  ..., -0.3714, -0.1012,  0.1389],
         [-0.0712, -0.0562, -0.0412,  ..., -0.3564, -0.1763,  0.0789],
         [-0.2663, -0.3114, -0.4014,  ..., -0.2363, -0.2063, -0.2663]],

        [[ 0.1835,  0.2120,  0.0129,  ...,  0.1124, -0.1862, -0.7123],
         [ 0.2262,  0.2973,  0.0413,  ...,  0.2404,  0.3257,  0.2546],
         [ 0.2262,  0.2120,  0.2120,  ...,  0.2831,  0.2404,  0.2120],
         ...,
         [-0.8688, -0.9256, -0.8972,  ..., -0.7977, -0.6555, -0.2289],
         [-0.6697, -0.7977, -0.6839,  ..., -0.8119, -0.7123, -0.3142],
         [-1.0252, -1.1105, -1.0821,  ..., -0.5559, -0.6555, -0.8119]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the cup thing in the upper left corner in this picture. ASSISTANT: cup thing in the upper left corner.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the glass top right corner of table in this picture? ASSISTANT: Here's the area for glass top right corner of table.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the top left corner item in this image. ASSISTANT: This view focuses on the top left corner item.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the glass container in this image. ASSISTANT: I've marked the glass container for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the cup with brown stuff in it in this image? ASSISTANT: Result: cup with brown stuff in it.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [1, 1, 1,  ..., 0, 0, 0],
         [1, 1, 1,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [1, 1, 1,  ..., 0, 0, 0],
         [1, 1, 1,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nPlease point out the cup thing in the upper left corner in this picture.', '<image>\nCould you isolate the glass top right corner of table in this picture?', '<image>\nPlease highlight the top left corner item in this image.', '<image>\nPlease highlight the glass container in this image.', '<image>\nCan you segment the cup with brown stuff in it in this image?'], ['cup thing in the upper left corner', 'glass top right corner of table', 'top left corner item', 'glass container', 'cup with brown stuff in it'], [{'cup thing in the upper left corner': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'glass top right corner of table': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'top left corner item': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0],
        [1, 1, 1,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'glass container': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'cup with brown stuff in it': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000033787.jpg', tensor([[[-0.9020, -0.8164, -0.7137,  ..., -0.5424, -0.5596, -0.5596],
         [-1.3987, -1.1760, -0.8507,  ..., -0.5424, -0.5596, -0.5596],
         [-2.0323, -1.6384, -1.0904,  ..., -0.5253, -0.5424, -0.5424],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.7927, -0.5476, -0.2675,  ...,  0.3277,  0.3102,  0.3102],
         [-1.2304, -0.9328, -0.5651,  ...,  0.3277,  0.3102,  0.3102],
         [-1.7906, -1.4580, -0.9853,  ...,  0.3452,  0.3277,  0.3277],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.3578, -0.0092,  0.3916,  ...,  1.4722,  1.4548,  1.4548],
         [-0.9156, -0.3753,  0.2871,  ...,  1.4897,  1.4722,  1.4722],
         [-1.6127, -0.8807,  0.0953,  ...,  1.5245,  1.5071,  1.5071],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.2515, -1.4857, -1.4711,  ..., -0.3032, -0.2886, -0.2886],
         [ 0.1201, -0.9602, -1.5149,  ..., -0.2886, -0.2740, -0.2886],
         [ 0.4705,  0.0179, -1.2959,  ..., -0.2740, -0.2886, -0.2740],
         ...,
         [-0.5952, -0.5368, -0.5076,  ..., -0.6536, -0.6098, -0.6098],
         [-0.5368, -0.5076, -0.5368,  ..., -0.6536, -0.6682, -0.6390],
         [-0.6390, -0.7266, -0.7266,  ..., -0.6828, -0.6828, -0.6682]],

        [[ 0.8743, -1.4069, -1.4069,  ...,  0.5591,  0.5591,  0.5441],
         [ 0.6792, -0.7016, -1.5120,  ...,  0.5591,  0.5741,  0.5591],
         [ 1.0393,  0.4540, -1.2268,  ...,  0.5741,  0.5741,  0.5741],
         ...,
         [-0.0262,  0.0638,  0.0939,  ..., -0.1913, -0.1313, -0.1463],
         [ 0.0038,  0.0188, -0.0262,  ..., -0.2063, -0.2063, -0.1913],
         [-0.2063, -0.3114, -0.3114,  ..., -0.2663, -0.2363, -0.2363]],

        [[ 1.3638, -1.1389, -1.1105,  ...,  1.5202,  1.5202,  1.5060],
         [ 1.1789, -0.3284, -1.3238,  ...,  1.5344,  1.5344,  1.5202],
         [ 1.6482,  1.0794, -0.9825,  ...,  1.5202,  1.5202,  1.5344],
         ...,
         [ 0.6812,  0.7523,  0.7950,  ...,  0.5106,  0.5106,  0.5248],
         [ 0.7097,  0.7239,  0.6955,  ...,  0.4821,  0.4679,  0.4679],
         [ 0.4537,  0.3542,  0.3684,  ...,  0.3968,  0.4537,  0.4537]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the main object in the image and where is it located? ASSISTANT: The main object in the image is a bench, which is located next to a tree and looks out onto a snow-covered field.</s>USER: What is the general setting and season of the image? ASSISTANT: The image is set in a park or a field during winter, with snow covering both the bench and the surrounding landscape.</s>USER: Is there any body of water such as a lake visible in the image? ASSISTANT: Yes, there is a snow-covered frozen lake visible in the image, and the bench is situated along its shore.</s>USER: How many trees can be seen next to the bench? ASSISTANT: There is one tree standing next to the bench in the image.</s>USER: What is the overall atmosphere of the image? Does it evoke a sense of solitude or liveliness? ASSISTANT: The overall atmosphere of the image is quite serene and tranquil, evoking a sense of solitude. The empty park bench, the snow-covered landscape, and the quiet stillness of the frozen lake create a peaceful, contemplative scene that would be ideal for someone seeking a moment of reflection or solitude. The absence of people in the scene, as well as the minimalistic composition encompassing just the bench, tree, and snow-covered field, further contribute to the sense of calm and solitude.</s>"], tensor([], size=(0, 427, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (683, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the main object in the image and where is it located? ASSISTANT: The main object in the image is a bench, which is located next to a tree and looks out onto a snow-covered field.</s>USER: What is the general setting and season of the image? ASSISTANT: The image is set in a park or a field during winter, with snow covering both the bench and the surrounding landscape.</s>USER: Is there any body of water such as a lake visible in the image? ASSISTANT: Yes, there is a snow-covered frozen lake visible in the image, and the bench is situated along its shore.</s>USER: How many trees can be seen next to the bench? ASSISTANT: There is one tree standing next to the bench in the image.</s>USER: What is the overall atmosphere of the image? Does it evoke a sense of solitude or liveliness? ASSISTANT: The overall atmosphere of the image is quite serene and tranquil, evoking a sense of solitude. The empty park bench, the snow-covered landscape, and the quiet stillness of the frozen lake create a peaceful, contemplative scene that would be ideal for someone seeking a moment of reflection or solitude. The absence of people in the scene, as well as the minimalistic composition encompassing just the bench, tree, and snow-covered field, further contribute to the sense of calm and solitude.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the main object in the image and where is it located? ASSISTANT: The main object in the image is a bench, which is located next to a tree and looks out onto a snow-covered field.</s>USER: What is the general setting and season of the image? ASSISTANT: The image is set in a park or a field during winter, with snow covering both the bench and the surrounding landscape.</s>USER: Is there any body of water such as a lake visible in the image? ASSISTANT: Yes, there is a snow-covered frozen lake visible in the image, and the bench is situated along its shore.</s>USER: How many trees can be seen next to the bench? ASSISTANT: There is one tree standing next to the bench in the image.</s>USER: What is the overall atmosphere of the image? Does it evoke a sense of solitude or liveliness? ASSISTANT: The overall atmosphere of the image is quite serene and tranquil, evoking a sense of solitude. The empty park bench, the snow-covered landscape, and the quiet stillness of the frozen lake create a peaceful, contemplative scene that would be ideal for someone seeking a moment of reflection or solitude. The absence of people in the scene, as well as the minimalistic composition encompassing just the bench, tree, and snow-covered field, further contribute to the sense of calm and solitude.</s>"], [{}], False, 'vqa')]
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/srWQsVZVnxp6wH8b2qCUCA.jpg', tensor([[[1.4098, 1.4098, 1.4098,  ..., 2.2489, 2.2489, 2.2489],
         [1.4098, 1.4098, 1.4098,  ..., 2.2489, 2.2489, 2.2489],
         [1.4098, 1.4098, 1.4098,  ..., 2.2489, 2.2489, 2.2489],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.0084, 2.0084, 2.0084,  ..., 2.4286, 2.4286, 2.4286],
         [2.0084, 2.0084, 2.0084,  ..., 2.4286, 2.4286, 2.4286],
         [2.0084, 2.0084, 2.0084,  ..., 2.4286, 2.4286, 2.4286],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.3437, 2.3437, 2.3437,  ..., 2.6400, 2.6400, 2.6400],
         [2.3437, 2.3437, 2.3437,  ..., 2.6400, 2.6400, 2.6400],
         [2.3437, 2.3437, 2.3437,  ..., 2.6400, 2.6400, 2.6400],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 9.5224e-01,  9.6684e-01,  9.6684e-01,  ...,  1.9303e+00,
           1.9303e+00,  1.9303e+00],
         [ 9.5224e-01,  9.6684e-01,  9.8144e-01,  ...,  1.9303e+00,
           1.9303e+00,  1.9303e+00],
         [ 9.3764e-01,  9.5224e-01,  9.9604e-01,  ...,  1.9303e+00,
           1.9303e+00,  1.9303e+00],
         ...,
         [-5.6599e-01, -5.6599e-01, -4.9300e-01,  ..., -4.7840e-01,
          -4.2001e-01, -4.2001e-01],
         [-5.5140e-01, -5.8059e-01, -5.8059e-01,  ..., -4.7840e-01,
          -4.4921e-01, -4.4921e-01],
         [-6.0979e-01, -6.0979e-01, -6.3899e-01,  ..., -4.0541e-01,
          -4.0541e-01, -4.0541e-01]],

        [[ 1.4746e+00,  1.5046e+00,  1.5346e+00,  ...,  2.0749e+00,
           2.0749e+00,  2.0749e+00],
         [ 1.4896e+00,  1.5196e+00,  1.5346e+00,  ...,  2.0749e+00,
           2.0749e+00,  2.0749e+00],
         [ 1.4896e+00,  1.5046e+00,  1.5046e+00,  ...,  2.0749e+00,
           2.0749e+00,  2.0749e+00],
         ...,
         [-2.2130e-01, -2.2130e-01, -1.6127e-01,  ..., -1.3126e-01,
          -8.6235e-02, -8.6235e-02],
         [-1.9129e-01, -2.3631e-01, -2.5132e-01,  ..., -1.1625e-01,
          -1.0124e-01, -1.1625e-01],
         [-2.5132e-01, -2.9634e-01, -3.1135e-01,  ..., -4.1212e-02,
          -4.1212e-02, -7.1227e-02]],

        [[ 1.7051e+00,  1.7051e+00,  1.7051e+00,  ...,  2.1459e+00,
           2.1459e+00,  2.1459e+00],
         [ 1.6909e+00,  1.7051e+00,  1.7193e+00,  ...,  2.1459e+00,
           2.1459e+00,  2.1459e+00],
         [ 1.6766e+00,  1.6909e+00,  1.7051e+00,  ...,  2.1459e+00,
           2.1459e+00,  2.1459e+00],
         ...,
         [-5.8213e-02, -2.9773e-02,  1.2887e-02,  ..., -1.5553e-02,
          -1.3329e-03,  1.2887e-02],
         [-2.9773e-02, -5.8213e-02, -7.2433e-02,  ..., -1.3329e-03,
           1.2887e-02, -1.5553e-02],
         [-8.6653e-02, -1.1509e-01, -1.2931e-01,  ...,  6.9767e-02,
           8.3987e-02,  4.1327e-02]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the lane marking - straight line from the provided image. ASSISTANT: The lane marking - straight line portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the building in this picture? ASSISTANT: Take a look at the building here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the traffic island from the provided image. ASSISTANT: You can see the traffic island in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the lane marking - other from the provided image. ASSISTANT: You can see the lane marking - other in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is curb in this image? Please give some explanation. ASSISTANT: This is the selected curb.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ...,  True,  True,  True],
         [False, False, False,  ...,  True,  True,  True],
         [False, False, False,  ...,  True,  True,  True]]]), tensor([[61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        [61, 61, 61,  ..., 61, 61, 61],
        ...,
        [21, 21, 21,  ...,  4,  4,  4],
        [21, 21, 21,  ...,  4,  4,  4],
        [21, 21, 21,  ...,  4,  4,  4]]), (768, 1024), ['<image>\nSegment the lane marking - straight line from the provided image.', '<image>\nCould you isolate the building in this picture?', '<image>\nSegment the traffic island from the provided image.', '<image>\nSegment the lane marking - other from the provided image.', '<image>\nWhat is curb in this image? Please give some explanation.'], ['lane marking - straight line', 'building', 'traffic island', 'lane marking - other', 'curb'], [{'lane marking - straight line': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'building': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'traffic island': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'lane marking - other': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'curb': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ...,  True,  True,  True],
        [False, False, False,  ...,  True,  True,  True],
        [False, False, False,  ...,  True,  True,  True]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000187566.jpg', tensor([[[-0.5424, -0.5596, -0.5938,  ..., -0.7137, -0.6109, -0.5253],
         [-0.8678, -0.8849, -0.8849,  ..., -0.7137, -0.6281, -0.5596],
         [-1.2617, -1.2617, -1.2445,  ..., -0.6965, -0.6452, -0.6109],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.3452,  0.3277,  0.3102,  ..., -1.1078, -0.2675,  0.3978],
         [-0.3375, -0.3375, -0.3375,  ..., -1.0903, -0.2675,  0.3803],
         [-1.1779, -1.1604, -1.1429,  ..., -1.0728, -0.2675,  0.3452],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.6640,  1.6814,  1.7163,  ..., -1.1596,  0.4439,  1.6988],
         [ 0.4614,  0.4788,  0.4962,  ..., -1.1421,  0.4439,  1.6814],
         [-1.0201, -1.0201, -1.0376,  ..., -1.1247,  0.4265,  1.6465],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.6536, -0.6536, -0.6244,  ...,  0.5435,  0.4121,  0.4851],
         [ 0.3829,  0.3537,  0.3245,  ..., -0.9456, -0.9748, -1.0477],
         [-0.3616, -0.4054, -0.3616,  ..., -1.4127, -1.4273, -1.4273],
         ...,
         [-1.5295, -1.5149, -1.5149,  ..., -1.4127, -0.5368,  0.6749],
         [-1.5441, -1.5149, -1.5295,  ..., -1.4127, -0.6244,  0.5289],
         [-1.3105, -1.3105, -1.3397,  ..., -1.3835, -1.0769,  0.3099]],

        [[-0.2963, -0.3114, -0.2663,  ...,  0.8893,  0.7392,  0.7542],
         [ 0.7242,  0.7092,  0.6942,  ..., -0.8816, -0.9117, -0.9867],
         [-0.2063, -0.2513, -0.2063,  ..., -1.3919, -1.3619, -1.3919],
         ...,
         [-1.4669, -1.4369, -1.4669,  ..., -1.4669, -1.1068, -0.5365],
         [-1.4820, -1.4519, -1.4519,  ..., -1.4820, -1.0617, -0.6265],
         [-1.3019, -1.3019, -1.3169,  ..., -1.4970, -1.3469, -0.7166]],

        [[ 0.1551,  0.1551,  0.1977,  ...,  1.1647,  0.9799,  0.9088],
         [ 0.8234,  0.8234,  0.7239,  ..., -0.8545, -0.8545, -0.9256],
         [-0.3426, -0.3995, -0.3995,  ..., -1.1247, -1.1389, -1.1532],
         ...,
         [-1.1532, -1.1389, -1.1532,  ..., -1.1674, -1.2669, -1.4660],
         [-1.1532, -1.1247, -1.1247,  ..., -1.2100, -1.2811, -1.4518],
         [-1.0963, -1.0963, -1.1105,  ..., -1.2527, -1.2811, -1.4376]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nHow many types of meat are on the grill? ASSISTANT: There are two types of meat on the grill: steak and hot dogs.</s>USER: How would you describe the hot dogs in the image? ASSISTANT: The hot dogs in the image are of a standard size and shape, cooking on the grill alongside the steak.</s>USER: What is happening to the steak and hot dogs in the image? ASSISTANT: The steak and hot dogs are being cooked on the grill, getting exposure to heat and likely developing a charred or grilled appearance and texture.</s>USER: What are some differences between cooking steak and hot dogs on a grill? ASSISTANT: There are several differences between cooking steak and hot dogs on a grill:\n\n1. Cooking time: Steaks typically require more time to cook on a grill compared to hot dogs. Steaks are thicker and have a higher mass, which means they need to cook longer to reach the desired internal temperature. In contrast, hot dogs are thinner and precooked, requiring less time on the grill to heat through and develop a grilled appearance.\n\n2. Temperature: Steaks are usually cooked at a higher temperature to develop a sear on the outside while not overcooking the inside. A hot grill sears the outside quickly, preserving the tenderness and juiciness of the meat inside. Hot dogs, on the other hand, can be cooked at a lower temperature since they are already precooked, and the primary goal is to heat them through and add some grill marks.\n\n3. Grilling techniques: For steak, it is important to flip the meat only once or twice during cooking to ensure an even sear and avoid overcooking. Over-flipping a steak may result in uneven cooking and a less desirable final result. With hot dogs, more frequent flipping is not as much of a concern, and they can be rotated more often to cook evenly and achieve a uniform grilled appearance.\n\n4. Doneness levels: Steaks can be cooked to various levels of doneness such as rare, medium-rare, medium, or well-done, which depends on the desired internal temperature and personal preference of the person eating the steak. Hot dogs, being precooked, only need to be heated to a safe eating temperature and don't require specific doneness levels.</s>"], tensor([], size=(0, 601, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (962, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nHow many types of meat are on the grill? ASSISTANT: There are two types of meat on the grill: steak and hot dogs.</s>USER: How would you describe the hot dogs in the image? ASSISTANT: The hot dogs in the image are of a standard size and shape, cooking on the grill alongside the steak.</s>USER: What is happening to the steak and hot dogs in the image? ASSISTANT: The steak and hot dogs are being cooked on the grill, getting exposure to heat and likely developing a charred or grilled appearance and texture.</s>USER: What are some differences between cooking steak and hot dogs on a grill? ASSISTANT: There are several differences between cooking steak and hot dogs on a grill:\n\n1. Cooking time: Steaks typically require more time to cook on a grill compared to hot dogs. Steaks are thicker and have a higher mass, which means they need to cook longer to reach the desired internal temperature. In contrast, hot dogs are thinner and precooked, requiring less time on the grill to heat through and develop a grilled appearance.\n\n2. Temperature: Steaks are usually cooked at a higher temperature to develop a sear on the outside while not overcooking the inside. A hot grill sears the outside quickly, preserving the tenderness and juiciness of the meat inside. Hot dogs, on the other hand, can be cooked at a lower temperature since they are already precooked, and the primary goal is to heat them through and add some grill marks.\n\n3. Grilling techniques: For steak, it is important to flip the meat only once or twice during cooking to ensure an even sear and avoid overcooking. Over-flipping a steak may result in uneven cooking and a less desirable final result. With hot dogs, more frequent flipping is not as much of a concern, and they can be rotated more often to cook evenly and achieve a uniform grilled appearance.\n\n4. Doneness levels: Steaks can be cooked to various levels of doneness such as rare, medium-rare, medium, or well-done, which depends on the desired internal temperature and personal preference of the person eating the steak. Hot dogs, being precooked, only need to be heated to a safe eating temperature and don't require specific doneness levels.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nHow many types of meat are on the grill? ASSISTANT: There are two types of meat on the grill: steak and hot dogs.</s>USER: How would you describe the hot dogs in the image? ASSISTANT: The hot dogs in the image are of a standard size and shape, cooking on the grill alongside the steak.</s>USER: What is happening to the steak and hot dogs in the image? ASSISTANT: The steak and hot dogs are being cooked on the grill, getting exposure to heat and likely developing a charred or grilled appearance and texture.</s>USER: What are some differences between cooking steak and hot dogs on a grill? ASSISTANT: There are several differences between cooking steak and hot dogs on a grill:\n\n1. Cooking time: Steaks typically require more time to cook on a grill compared to hot dogs. Steaks are thicker and have a higher mass, which means they need to cook longer to reach the desired internal temperature. In contrast, hot dogs are thinner and precooked, requiring less time on the grill to heat through and develop a grilled appearance.\n\n2. Temperature: Steaks are usually cooked at a higher temperature to develop a sear on the outside while not overcooking the inside. A hot grill sears the outside quickly, preserving the tenderness and juiciness of the meat inside. Hot dogs, on the other hand, can be cooked at a lower temperature since they are already precooked, and the primary goal is to heat them through and add some grill marks.\n\n3. Grilling techniques: For steak, it is important to flip the meat only once or twice during cooking to ensure an even sear and avoid overcooking. Over-flipping a steak may result in uneven cooking and a less desirable final result. With hot dogs, more frequent flipping is not as much of a concern, and they can be rotated more often to cook evenly and achieve a uniform grilled appearance.\n\n4. Doneness levels: Steaks can be cooked to various levels of doneness such as rare, medium-rare, medium, or well-done, which depends on the desired internal temperature and personal preference of the person eating the steak. Hot dogs, being precooked, only need to be heated to a safe eating temperature and don't require specific doneness levels.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000203098.jpg', tensor([[[-0.5424, -0.5082, -0.4397,  ...,  0.4679,  0.1768,  0.0398],
         [-0.3541, -0.3712, -0.4054,  ...,  0.3823,  0.1254,  0.0227],
         [ 0.0569, -0.0629, -0.3198,  ...,  0.1768,  0.0398, -0.0287],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.0553, -1.0378, -1.0028,  ..., -0.5126, -0.7927, -0.9328],
         [-0.8277, -0.8627, -0.9328,  ..., -0.5651, -0.7927, -0.9153],
         [-0.3550, -0.4951, -0.7927,  ..., -0.6702, -0.8102, -0.8803],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.0376, -1.0376, -1.0550,  ..., -0.8981, -1.1247, -1.2467],
         [-0.8633, -0.9156, -1.0201,  ..., -0.9330, -1.1247, -1.2119],
         [-0.4798, -0.6367, -0.9504,  ..., -1.0201, -1.1073, -1.1596],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.1864, -0.4346, -0.3762,  ..., -0.7850, -0.6244, -0.1718],
         [-0.6828, -0.6390, -0.5806,  ..., -0.0696, -0.7412, -0.5222],
         [-0.1718,  0.0325,  0.0179,  ..., -0.0113, -0.8142, -0.8726],
         ...,
         [ 0.0179,  0.5581,  0.3683,  ..., -0.5660, -0.9164,  0.0763],
         [-0.6974, -0.5806,  0.0617,  ...,  0.6019, -0.5660,  0.1055],
         [-0.4492, -0.2302,  0.4997,  ...,  0.6895,  0.5873,  0.5873]],

        [[-0.7016, -0.9267, -0.8666,  ..., -1.3019, -1.0167, -0.7766],
         [-1.0617, -1.0017, -0.9567,  ..., -0.6415, -1.1368, -0.9717],
         [-0.9417, -0.7016, -0.6715,  ..., -0.7616, -1.3019, -1.2718],
         ...,
         [-0.4914, -0.1012, -0.2963,  ..., -1.1218, -1.3019, -0.5215],
         [-1.2268, -1.0167, -0.7916,  ..., -0.0562, -1.1068, -0.5065],
         [-1.0017, -0.7166, -0.3114,  ..., -0.0862, -0.0562, -0.0712]],

        [[-0.8119, -0.9967, -0.9399,  ..., -1.2385, -0.9399, -0.8688],
         [-1.0110, -0.9399, -0.8688,  ..., -0.8688, -1.1105, -0.9825],
         [-1.1247, -0.9256, -0.8545,  ..., -0.9825, -1.2527, -1.1105],
         ...,
         [-0.7408, -0.4706, -0.5844,  ..., -1.1389, -1.1958, -0.7408],
         [-1.1958, -1.0678, -0.8119,  ..., -0.3000, -1.2243, -0.8545],
         [-1.0252, -0.8403, -0.5986,  ..., -0.3711, -0.4137, -0.4279]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is this photo about'? ASSISTANT: The image displays a variety of food items arranged on a woven surface. There are four bowls, each filled with different types of foods. The first bowl is located on the left side, the second bowl is situated in the bottom left corner, the third bowl is placed towards the right side of the scene, and the fourth bowl can be found at the top right corner.\n\nAside from the bowls, there are a couple of slices of pizza nearby, one positioned towards the middle of the scene and the other placed near the bottom area. Also, several carrots are scattered throughout the scene, adding color and variety to the food presentation.</s>"], tensor([], size=(0, 478, 500)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (979, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is this photo about'? ASSISTANT: The image displays a variety of food items arranged on a woven surface. There are four bowls, each filled with different types of foods. The first bowl is located on the left side, the second bowl is situated in the bottom left corner, the third bowl is placed towards the right side of the scene, and the fourth bowl can be found at the top right corner.\n\nAside from the bowls, there are a couple of slices of pizza nearby, one positioned towards the middle of the scene and the other placed near the bottom area. Also, several carrots are scattered throughout the scene, adding color and variety to the food presentation.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is this photo about'? ASSISTANT: The image displays a variety of food items arranged on a woven surface. There are four bowls, each filled with different types of foods. The first bowl is located on the left side, the second bowl is situated in the bottom left corner, the third bowl is placed towards the right side of the scene, and the fourth bowl can be found at the top right corner.\n\nAside from the bowls, there are a couple of slices of pizza nearby, one positioned towards the middle of the scene and the other placed near the bottom area. Also, several carrots are scattered throughout the scene, adding color and variety to the food presentation.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2010_003612.jpg', tensor([[[0.0056, 0.0569, 0.1426,  ..., 0.3309, 0.3138, 0.2967],
         [0.0398, 0.0741, 0.1426,  ..., 0.3309, 0.3138, 0.2967],
         [0.1254, 0.1254, 0.1426,  ..., 0.3138, 0.3138, 0.3138],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.1352, 0.1877, 0.2927,  ..., 0.4503, 0.4328, 0.4153],
         [0.1877, 0.2227, 0.2927,  ..., 0.4503, 0.4328, 0.4153],
         [0.2752, 0.2752, 0.2927,  ..., 0.4328, 0.4328, 0.4328],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.3219, 0.3568, 0.4265,  ..., 0.6182, 0.6008, 0.6008],
         [0.3568, 0.3742, 0.4265,  ..., 0.6182, 0.6008, 0.6008],
         [0.4265, 0.4265, 0.4091,  ..., 0.6008, 0.5834, 0.5834],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 0.4121,  0.3975,  0.4267,  ...,  0.0179,  0.4851,  0.5727],
         [ 0.4267,  0.4267,  0.4413,  ...,  0.0179,  0.5143,  0.5873],
         [ 0.4267,  0.4413,  0.4413,  ...,  0.0471,  0.4851,  0.5727],
         ...,
         [-1.3981, -1.3251, -1.3105,  ..., -1.2667, -1.3105, -1.2667],
         [-1.3981, -1.3835, -1.3689,  ..., -1.2521, -1.3105, -1.3397],
         [-1.3835, -1.3981, -1.3981,  ..., -1.2813, -1.2229, -1.2959]],

        [[ 0.5141,  0.4991,  0.5291,  ...,  0.0638,  0.5741,  0.7242],
         [ 0.5291,  0.5291,  0.5291,  ...,  0.0939,  0.6191,  0.6942],
         [ 0.5291,  0.5141,  0.5291,  ...,  0.1389,  0.6041,  0.6942],
         ...,
         [-1.1668, -1.0918, -1.0617,  ..., -0.9417, -1.0017, -0.9717],
         [-1.1218, -1.1068, -1.1218,  ..., -0.9267, -1.0317, -1.0467],
         [-1.0918, -1.1518, -1.1968,  ..., -0.9417, -0.9267, -1.0017]],

        [[ 0.6812,  0.6812,  0.6812,  ...,  0.1551,  0.6955,  0.8519],
         [ 0.6812,  0.6670,  0.6528,  ...,  0.1977,  0.7239,  0.8234],
         [ 0.7097,  0.6812,  0.6812,  ...,  0.2404,  0.6812,  0.7808],
         ...,
         [-0.7123, -0.6270, -0.5133,  ..., -0.4279, -0.4990, -0.5275],
         [-0.6697, -0.6412, -0.5986,  ..., -0.3711, -0.4706, -0.5417],
         [-0.5844, -0.6697, -0.6981,  ..., -0.3711, -0.3568, -0.4564]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the car wheel shown in this image. ASSISTANT: car wheel.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the person arm shown in this image. ASSISTANT: Result: person arm.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the the leg of the person from the image below? ASSISTANT: Here is the the leg of the person you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the the leg of the person in this picture. ASSISTANT: This view focuses on the the leg of the person.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the window of the car in this image? Please give some explanation. ASSISTANT: This view focuses on the the window of the car.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (680, 1024), ['<image>\nDisplay a segmentation mask for the car wheel shown in this image.', '<image>\nDisplay a segmentation mask for the person arm shown in this image.', '<image>\nWould you please extract the the leg of the person from the image below?', '<image>\nPlease point out the the leg of the person in this picture.', '<image>\nWhat is the window of the car in this image? Please give some explanation.'], ['car wheel', 'person arm', 'the leg of the person', 'the leg of the person', 'the window of the car'], [{'car wheel': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person arm': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the leg of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the leg of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the window of the car': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000325229.jpg', tensor([[[ 0.1597, -0.0287, -0.4397,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.1768,  0.0056, -0.3883,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.2282,  0.0569, -0.2856,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 2.0092,  1.9920,  1.9578,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.7009,  1.6153,  1.4612,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.5468,  1.4440,  1.2214,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.3627,  0.1877, -0.1975,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.3803,  0.2052, -0.1625,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.3978,  0.2577, -0.0924,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 1.9384,  1.9559,  1.9909,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.5707,  1.5357,  1.4307,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.3957,  1.3256,  1.1681,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.2871,  0.0431, -0.4624,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.3219,  0.0779, -0.4275,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.3742,  0.1302, -0.3753,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [ 2.0300,  2.0300,  2.0474,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.6988,  1.6465,  1.5245,  ...,  0.0000,  0.0000,  0.0000],
         [ 1.5420,  1.4548,  1.2805,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.4273, -1.3251, -1.1061,  ...,  0.3829,  0.1785,  0.1785],
         [-1.5149, -1.1061, -1.1937,  ...,  0.9960,  0.9376,  1.3756],
         [-1.3397, -1.0623, -1.2375,  ...,  0.7771,  0.7625,  0.6895],
         ...,
         [ 0.4705,  0.7187,  0.6895,  ...,  0.2077,  0.0179, -0.1426],
         [ 0.5435,  0.7625,  0.9376,  ..., -0.0988, -0.1134,  0.4559],
         [ 1.0544,  0.8355,  1.2004,  ..., -0.2302, -0.3762, -0.2740]],

        [[-1.0467, -0.9567, -0.8216,  ...,  0.7992,  0.6191,  0.6792],
         [-1.0017, -0.7916, -1.0467,  ...,  1.2194,  1.1594,  1.6547],
         [-0.9567, -0.7166, -1.1668,  ...,  1.2495,  1.2194,  1.1894],
         ...,
         [ 0.8743,  1.0243,  1.0393,  ...,  0.2890,  0.0638, -0.0712],
         [ 1.0393,  0.9943,  0.8292,  ..., -0.1163, -0.1763,  0.4090],
         [ 1.0093,  0.6942,  1.0694,  ..., -0.1763, -0.2813, -0.1613]],

        [[-1.0963, -0.9541, -0.8261,  ...,  1.2358,  1.1932,  1.0794],
         [-1.0963, -0.9256, -0.9825,  ...,  1.6055,  1.6055,  1.8615],
         [-0.9256, -0.9399, -1.1389,  ...,  1.4207,  1.4776,  1.4065],
         ...,
         [ 0.0413,  0.1266,  0.2831,  ...,  0.3684,  0.1977,  0.0129],
         [ 0.1409,  0.2688,  0.3542,  ...,  0.1124, -0.0156,  0.4395],
         [ 0.3399,  0.2404,  0.5675,  ..., -0.0582, -0.2715, -0.0582]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the garaif in back in this image? ASSISTANT: I've marked the garaif in back for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is right giraffe in this image? Please elaborate your answer and explain why. ASSISTANT: Result: right giraffe.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the giraffe on the left in this picture. ASSISTANT: Below you can see the giraffe on the left.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the one on left in this picture? ASSISTANT: Take a look at the one on left here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the front giraffe in this image. ASSISTANT: front giraffe.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (1024, 768), ['<image>\nCan you segment the garaif in back in this image?', '<image>\nWhat is right giraffe in this image? Please elaborate your answer and explain why.', '<image>\nPlease point out the giraffe on the left in this picture.', '<image>\nCould you identify and segment out the one on left in this picture?', '<image>\nPlease highlight the front giraffe in this image.'], ['garaif in back', 'right giraffe', 'giraffe on the left', 'one on left', 'front giraffe'], [{'garaif in back': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'right giraffe': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'giraffe on the left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'one on left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'front giraffe': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/oW8AgQnqYuL7lzu8D5YabA.jpg', tensor([[[-1.7069, -1.3473, -0.7650,  ..., -0.9192, -1.0219, -1.0219],
         [-0.9363, -0.7993, -0.5082,  ..., -0.9705, -1.1075, -1.1589],
         [-0.6623, -0.1314, -0.7308,  ..., -1.0562, -1.1589, -1.2445],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.1078, -0.9328,  0.4153,  ...,  0.6078,  0.5728,  0.5553],
         [-0.1800,  0.0651,  0.8529,  ...,  0.5728,  0.6078,  0.5728],
         [ 0.6254,  1.0980,  0.2577,  ...,  0.5553,  0.6254,  0.5903],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.6705,  0.7402,  1.4200,  ...,  1.9603,  1.9080,  1.8905],
         [ 1.2631,  1.2980,  1.6117,  ...,  1.9254,  1.8905,  1.8557],
         [ 1.5594,  1.9080,  0.9842,  ...,  1.8905,  1.8905,  1.8383],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.3245,  0.3391,  0.3391,  ..., -0.3178, -0.3324, -0.3762],
         [ 0.3245,  0.3245,  0.2807,  ..., -0.3324, -0.3178, -0.3470],
         [ 0.2953,  0.2515,  0.2223,  ..., -0.3178, -0.3178, -0.3324],
         ...,
         [-0.6098, -0.9018, -0.9018,  ..., -0.2740,  0.3975,  0.4997],
         [-0.9602, -0.9456, -0.7704,  ..., -0.3616,  0.2223,  0.5581],
         [-0.8726, -0.8434, -0.7850,  ..., -0.4054,  0.1055,  0.6749]],

        [[ 1.2344,  1.2495,  1.2495,  ...,  0.7692,  0.7542,  0.7392],
         [ 1.2344,  1.2344,  1.2495,  ...,  0.7542,  0.7692,  0.7392],
         [ 1.2495,  1.2194,  1.2194,  ...,  0.7542,  0.7692,  0.7542],
         ...,
         [-0.2663, -0.5065, -0.5815,  ..., -0.6865,  0.4991,  0.7242],
         [-0.6265, -0.5815, -0.4314,  ..., -0.8066,  0.2139,  0.7692],
         [-0.5365, -0.4914, -0.4464,  ..., -0.8516, -0.1313,  0.7542]],

        [[ 2.0037,  2.0179,  2.0179,  ...,  1.6909,  1.6766,  1.6624],
         [ 1.9753,  2.0037,  1.9895,  ...,  1.6624,  1.6624,  1.6482],
         [ 2.0037,  1.9753,  1.9753,  ...,  1.6624,  1.6624,  1.6482],
         ...,
         [-0.0156, -0.2289, -0.2715,  ..., -0.6270,  0.5248,  0.7666],
         [-0.3568, -0.3000, -0.1862,  ..., -0.7408,  0.2546,  0.8234],
         [-0.2715, -0.2289, -0.1862,  ..., -0.8261, -0.1720,  0.7239]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the lane marking - ambiguous in this image. ASSISTANT: I've marked the lane marking - ambiguous for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the traffic cone in this image. ASSISTANT: Here's the area for traffic cone.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the other vehicle in this picture? ASSISTANT: The other vehicle portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is terrain in this image? Please elaborate your answer and explain why. ASSISTANT: Here's the area for terrain.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the lane marking - straight line in this picture. ASSISTANT: Here's where the lane marking - straight line appears in the image.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[ 64,  64,  64,  ..., 123, 123,  61],
        [ 64,  64,  64,  ...,  61,  61,  61],
        [ 64,  64,  64,  ...,  61,  61,  61],
        ...,
        [ 21,  21,  21,  ...,  21,  21,  21],
        [ 21,  21,  21,  ...,  21,  21,  21],
        [ 21,  21,  21,  ...,  21,  21,  21]]), (768, 1024), ['<image>\nPlease highlight the lane marking - ambiguous in this image.', '<image>\nPlease highlight the traffic cone in this image.', '<image>\nCould you isolate the other vehicle in this picture?', '<image>\nWhat is terrain in this image? Please elaborate your answer and explain why.', '<image>\nPlease point out the lane marking - straight line in this picture.'], ['lane marking - ambiguous', 'traffic cone', 'other vehicle', 'terrain', 'lane marking - straight line'], [{'lane marking - ambiguous': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'traffic cone': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'other vehicle': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'terrain': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'lane marking - straight line': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg')]
>> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/reason_seg/ReasonSeg/train/3336271467_ae342dd581_o.jpg', tensor([[[ 0.9817,  0.9988,  0.9646,  ...,  0.1597,  0.1939,  0.1939],
         [ 0.9474,  0.9646,  0.9646,  ...,  0.1597,  0.1768,  0.1939],
         [ 0.9817,  0.9817,  0.9817,  ...,  0.1254,  0.1426,  0.1597],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.1331,  1.1331,  1.1331,  ...,  0.4153,  0.3978,  0.4153],
         [ 1.1506,  1.1331,  1.1506,  ...,  0.4153,  0.3978,  0.4153],
         [ 1.1506,  1.1506,  1.1856,  ...,  0.4328,  0.4328,  0.4503],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.0365,  1.0191,  1.0191,  ..., -0.2707, -0.2358, -0.2358],
         [ 1.0714,  1.0539,  1.0539,  ..., -0.2358, -0.2184, -0.2184],
         [ 1.0714,  1.0714,  1.0714,  ..., -0.2358, -0.2184, -0.2184],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.9668,  0.9230,  0.8647,  ...,  0.1201,  0.2515,  0.3975],
         [ 0.9522,  0.9084,  0.8501,  ...,  0.1201,  0.2515,  0.4121],
         [ 0.9376,  0.9084,  0.8501,  ...,  0.1201,  0.2515,  0.4267],
         ...,
         [-1.5587, -1.5295, -1.4419,  ...,  0.0179,  0.0033,  0.0179],
         [-1.5733, -1.5003, -1.3981,  ...,  0.0033,  0.0179,  0.0179],
         [-1.4711, -1.5003, -1.5003,  ...,  0.0179,  0.0325,  0.0179]],

        [[ 1.4145,  1.3545,  1.3095,  ...,  0.5291,  0.6191,  0.7242],
         [ 1.3995,  1.3695,  1.3245,  ...,  0.5291,  0.6491,  0.7542],
         [ 1.3845,  1.3545,  1.3095,  ...,  0.5441,  0.6642,  0.7842],
         ...,
         [-1.5420, -1.5120, -1.4669,  ...,  0.0488,  0.0038, -0.0712],
         [-1.5420, -1.4970, -1.4369,  ...,  0.0488,  0.0188, -0.0412],
         [-1.5420, -1.5570, -1.5120,  ...,  0.0338,  0.0038, -0.0262]],

        [[ 0.6101,  0.5675,  0.4821,  ..., -0.1009,  0.0555,  0.2404],
         [ 0.5959,  0.5675,  0.4964,  ..., -0.0724,  0.0698,  0.2688],
         [ 0.5675,  0.5248,  0.4821,  ..., -0.0440,  0.0840,  0.2546],
         ...,
         [-1.2527, -1.2243, -1.1389,  ..., -0.0867, -0.1720, -0.2004],
         [-1.2811, -1.1958, -1.1247,  ..., -0.0867, -0.1578, -0.2146],
         [-1.2385, -1.2527, -1.2243,  ..., -0.0867, -0.1293, -0.2004]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the unique physical feature of ostriches that sets them apart from other birds in the picture?. Please respond with segmentation mask. ASSISTANT: The unique physical attribute that sets this apart from other birds is found in its impressive size and long legs.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nOstriches are known for having a distinct physical feature that sets them apart from other birds. What part of the ostrich in the picture is unique and different from other birds?. Please respond with segmentation mask. ASSISTANT: The unique physical attribute that sets these apart from other birds is this.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat unique physical attribute of ostriches sets them apart from other birds, specifically in the picture provided?. Please respond with segmentation mask. ASSISTANT: The unique physical attribute that sets them apart from other birds is this.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat specific physical feature of the ostrich in the picture is unique and distinguishes it from other birds?. Please elaborate your answer and explain why. ASSISTANT: This unique physical attribute that sets them apart from other birds is their long legs.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat part of the ostrich present in the picture makes them unique and distinguishes them from other birds?. Please elaborate your answer and explain why. ASSISTANT: This distinct physical feature, notable for its size, sets it apart from other birds.</s>"], tensor([[[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]],

        [[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]],

        [[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]],

        [[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]],

        [[0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         ...,
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.],
         [0., 0., 0.,  ..., 0., 0., 0.]]]), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (683, 1024), ['<image>\nWhat is the unique physical feature of ostriches that sets them apart from other birds in the picture?. Please respond with segmentation mask.', '<image>\nOstriches are known for having a distinct physical feature that sets them apart from other birds. What part of the ostrich in the picture is unique and different from other birds?. Please respond with segmentation mask.', '<image>\nWhat unique physical attribute of ostriches sets them apart from other birds, specifically in the picture provided?. Please respond with segmentation mask.', '<image>\nWhat specific physical feature of the ostrich in the picture is unique and distinguishes it from other birds?. Please elaborate your answer and explain why.', '<image>\nWhat part of the ostrich present in the picture makes them unique and distinguishes them from other birds?. Please elaborate your answer and explain why.'], ['The unique physical attribute that sets this apart from other birds is found in its impressive size and long legs.', 'The unique physical attribute that sets these apart from other birds is this.', 'The unique physical attribute that sets them apart from other birds is this.', 'This unique physical attribute that sets them apart from other birds is their long legs.', 'This distinct physical feature, notable for its size, sets it apart from other birds.'], [{'What is the unique physical feature of ostriches that sets them apart from other birds in the picture?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}, {'Ostriches are known for having a distinct physical feature that sets them apart from other birds. What part of the ostrich in the picture is unique and different from other birds?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}, {'What unique physical attribute of ostriches sets them apart from other birds, specifically in the picture provided?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}, {'What specific physical feature of the ostrich in the picture is unique and distinguishes it from other birds?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}, {'What part of the ostrich present in the picture makes them unique and distinguishes them from other birds?': tensor([[0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        ...,
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.],
        [0., 0., 0.,  ..., 0., 0., 0.]])}], False, 'reason_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000018683.jpg', tensor([[[2.2489, 2.2489, 2.2489,  ..., 2.2489, 2.2489, 2.2489],
         [2.2489, 2.2489, 2.2489,  ..., 2.2489, 2.2489, 2.2489],
         [2.2489, 2.2489, 2.2489,  ..., 2.2489, 2.2489, 2.2489],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.4286, 2.4286, 2.4286,  ..., 2.4286, 2.4286, 2.4286],
         [2.4286, 2.4286, 2.4286,  ..., 2.4286, 2.4286, 2.4286],
         [2.4286, 2.4286, 2.4286,  ..., 2.4286, 2.4286, 2.4286],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.6400, 2.6400, 2.6400,  ..., 2.6400, 2.6400, 2.6400],
         [2.6400, 2.6400, 2.6400,  ..., 2.6400, 2.6400, 2.6400],
         [2.6400, 2.6400, 2.6400,  ..., 2.6400, 2.6400, 2.6400],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[1.9303, 1.9303, 1.9303,  ..., 1.9303, 1.9303, 1.9303],
         [1.9303, 1.9303, 1.9303,  ..., 1.9303, 1.9303, 1.9303],
         [1.9303, 1.9303, 1.9303,  ..., 1.9303, 1.9303, 1.9303],
         ...,
         [1.9303, 1.9303, 1.9303,  ..., 1.9303, 1.9303, 1.9303],
         [1.9303, 1.9303, 1.9303,  ..., 1.9303, 1.9303, 1.9303],
         [1.9303, 1.9303, 1.9303,  ..., 1.9303, 1.9303, 1.9303]],

        [[2.0749, 2.0749, 2.0749,  ..., 2.0749, 2.0749, 2.0749],
         [2.0749, 2.0749, 2.0749,  ..., 2.0749, 2.0749, 2.0749],
         [2.0749, 2.0749, 2.0749,  ..., 2.0749, 2.0749, 2.0749],
         ...,
         [2.0749, 2.0749, 2.0749,  ..., 2.0749, 2.0749, 2.0749],
         [2.0749, 2.0749, 2.0749,  ..., 2.0749, 2.0749, 2.0749],
         [2.0749, 2.0749, 2.0749,  ..., 2.0749, 2.0749, 2.0749]],

        [[2.1459, 2.1459, 2.1459,  ..., 2.1459, 2.1459, 2.1459],
         [2.1459, 2.1459, 2.1459,  ..., 2.1459, 2.1459, 2.1459],
         [2.1459, 2.1459, 2.1459,  ..., 2.1459, 2.1459, 2.1459],
         ...,
         [2.1459, 2.1459, 2.1459,  ..., 2.1459, 2.1459, 2.1459],
         [2.1459, 2.1459, 2.1459,  ..., 2.1459, 2.1459, 2.1459],
         [2.1459, 2.1459, 2.1459,  ..., 2.1459, 2.1459, 2.1459]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the girl holding phone left picture from the image below? ASSISTANT: Here is the girl holding phone left picture you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the girl right in this picture? ASSISTANT: I've marked the girl right for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the girl right in this picture? ASSISTANT: The girl right is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the pink in this picture. ASSISTANT: The pink portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the left girl shown in this image. ASSISTANT: Here is the left girl you asked about.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (790, 1024), ['<image>\nWould you please extract the girl holding phone left picture from the image below?', '<image>\nCould you identify and segment out the girl right in this picture?', '<image>\nCould you isolate the girl right in this picture?', '<image>\nPlease point out the pink in this picture.', '<image>\nDisplay a segmentation mask for the left girl shown in this image.'], ['girl holding phone left picture', 'girl right', 'girl right', 'pink', 'left girl'], [{'girl holding phone left picture': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'girl right': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'girl right': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'pink': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'left girl': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/ade20k/images/training/ADE_train_00005891.jpg', tensor([[[-1.5357, -1.5357, -1.5357,  ..., -1.9980, -1.9980, -1.9980],
         [-1.5357, -1.5357, -1.5357,  ..., -1.9980, -1.9980, -1.9980],
         [-1.5357, -1.5357, -1.5357,  ..., -1.9980, -1.9980, -1.9980],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.4055, -1.4055, -1.4055,  ..., -1.9132, -1.9132, -1.9132],
         [-1.4055, -1.4055, -1.4055,  ..., -1.9132, -1.9132, -1.9132],
         [-1.4055, -1.4055, -1.4055,  ..., -1.9132, -1.9132, -1.9132],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.4210, -1.4210, -1.4210,  ..., -1.6824, -1.6824, -1.6824],
         [-1.4210, -1.4210, -1.4210,  ..., -1.6824, -1.6824, -1.6824],
         [-1.4210, -1.4210, -1.4210,  ..., -1.6824, -1.6824, -1.6824],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.0039, -1.0185, -1.0185,  ..., -1.6463, -1.6901, -1.6463],
         [-1.0039, -1.0185, -1.0185,  ..., -1.6755, -1.6755, -1.6025],
         [-1.0039, -1.0185, -1.0185,  ..., -1.6317, -1.6171, -1.5879],
         ...,
         [-0.6828, -0.4492, -0.5514,  ...,  0.0325,  0.0909,  0.3391],
         [-0.7120, -0.6244, -0.4346,  ...,  0.0179,  0.1493,  0.3537],
         [-0.5660, -0.5222, -0.3762,  ..., -0.0113,  0.2223,  0.3829]],

        [[-0.9117, -0.9267, -0.9267,  ..., -1.5420, -1.5570, -1.5270],
         [-0.9117, -0.9267, -0.9267,  ..., -1.5870, -1.5720, -1.4970],
         [-0.9117, -0.9267, -0.9267,  ..., -1.5570, -1.5270, -1.4970],
         ...,
         [-0.7016, -0.4914, -0.6115,  ...,  0.2740,  0.2439,  0.4390],
         [-0.7466, -0.6715, -0.4914,  ...,  0.2289,  0.2740,  0.4390],
         [-0.5965, -0.5665, -0.4314,  ...,  0.1989,  0.3340,  0.4540]],

        [[-0.9967, -1.0110, -1.0110,  ..., -1.2811, -1.3238, -1.2811],
         [-0.9967, -1.0110, -1.0110,  ..., -1.3380, -1.3238, -1.2527],
         [-0.9967, -1.0110, -1.0110,  ..., -1.2954, -1.2811, -1.2527],
         ...,
         [-0.5844, -0.3568, -0.4706,  ...,  0.5390,  0.4679,  0.6101],
         [-0.6270, -0.5559, -0.3853,  ...,  0.4964,  0.4679,  0.5817],
         [-0.4990, -0.4706, -0.3284,  ...,  0.4537,  0.5106,  0.5959]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is wall in this image? Please elaborate your answer and explain why. ASSISTANT: Here is the wall you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the floor in this image? ASSISTANT: The floor portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the seat in this image. ASSISTANT: Here is the seat you asked about.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is crt screen in this image? Please give some explanation. ASSISTANT: The crt screen is presented here.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False,  True, False,  ..., False, False, False],
         [False,  True,  True,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[255, 255, 255,  ..., 255, 255, 255],
        [255,   0, 255,  ..., 255, 255, 255],
        [255,   0,   0,  ..., 255, 255, 255],
        ...,
        [ 31,  31,  31,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255]]), (768, 1024), ['<image>\nWhat is wall in this image? Please elaborate your answer and explain why.', '<image>\nCan you segment the floor in this image?', '<image>\nPlease segment the seat in this image.', '<image>\nWhat is crt screen in this image? Please give some explanation.'], ['wall', 'floor', 'seat', 'crt screen'], [{'wall': tensor([[False, False, False,  ..., False, False, False],
        [False,  True, False,  ..., False, False, False],
        [False,  True,  True,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'floor': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'seat': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'crt screen': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000163230.jpg', tensor([[[-1.3302, -1.3302, -1.3473,  ...,  0.0000,  0.0000,  0.0000],
         [-1.3302, -1.3302, -1.3473,  ...,  0.0000,  0.0000,  0.0000],
         [-1.3473, -1.3473, -1.3473,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-1.0048, -0.9877, -0.9705,  ...,  0.0000,  0.0000,  0.0000],
         [-0.9363, -0.9363, -0.9192,  ...,  0.0000,  0.0000,  0.0000],
         [-0.8849, -0.8849, -0.8678,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.7927, -0.7927, -0.8102,  ...,  0.0000,  0.0000,  0.0000],
         [-0.7927, -0.7927, -0.8102,  ...,  0.0000,  0.0000,  0.0000],
         [-0.8102, -0.8102, -0.8102,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.8803, -0.8627, -0.8452,  ...,  0.0000,  0.0000,  0.0000],
         [-0.8277, -0.8277, -0.8102,  ...,  0.0000,  0.0000,  0.0000],
         [-0.7927, -0.7927, -0.7752,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.0267, -0.0267, -0.0441,  ...,  0.0000,  0.0000,  0.0000],
         [-0.0441, -0.0441, -0.0441,  ...,  0.0000,  0.0000,  0.0000],
         [-0.0615, -0.0615, -0.0441,  ...,  0.0000,  0.0000,  0.0000],
         ...,
         [-0.3055, -0.2881, -0.2707,  ...,  0.0000,  0.0000,  0.0000],
         [-0.2707, -0.2707, -0.2532,  ...,  0.0000,  0.0000,  0.0000],
         [-0.2532, -0.2532, -0.2358,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.9164, -0.9310, -0.9164,  ..., -0.4784, -0.4784, -0.5076],
         [-0.9164, -0.9456, -0.9456,  ..., -0.4638, -0.4638, -0.4930],
         [-0.9018, -0.9310, -0.9164,  ..., -0.4638, -0.4638, -0.4638],
         ...,
         [ 1.5946,  1.6092,  1.6092,  ...,  1.5362,  1.5654,  1.5508],
         [ 1.5946,  1.5800,  1.5946,  ...,  1.5508,  1.5216,  1.5216],
         [ 1.5800,  1.5508,  1.6238,  ...,  1.4340,  1.4924,  1.4778]],

        [[-0.3264, -0.3114, -0.3264,  ...,  0.0488,  0.0638,  0.0488],
         [-0.3264, -0.3264, -0.3264,  ...,  0.0638,  0.0789,  0.0638],
         [-0.3114, -0.3114, -0.2963,  ...,  0.0789,  0.0638,  0.0638],
         ...,
         [ 1.7297,  1.7447,  1.7447,  ...,  1.6547,  1.6847,  1.6697],
         [ 1.7297,  1.7147,  1.7297,  ...,  1.6547,  1.6247,  1.6247],
         [ 1.7147,  1.6847,  1.7597,  ...,  1.5496,  1.6096,  1.5796]],

        [[ 0.3826,  0.3826,  0.3826,  ...,  0.7666,  0.7808,  0.7523],
         [ 0.3826,  0.3826,  0.3826,  ...,  0.7808,  0.7950,  0.7808],
         [ 0.3968,  0.3968,  0.4110,  ...,  0.7808,  0.7666,  0.7666],
         ...,
         [ 1.7904,  1.8046,  1.7904,  ...,  1.7193,  1.7477,  1.7335],
         [ 1.7904,  1.7762,  1.7904,  ...,  1.7477,  1.7193,  1.7193],
         [ 1.7762,  1.7477,  1.8188,  ...,  1.6482,  1.7051,  1.6909]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is mountain in this image? Please give some explanation. ASSISTANT: Below you can see the mountain.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the snow in this image. ASSISTANT: snow.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the snowboard shown in this image. ASSISTANT: Sure — this shows the snowboard.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the tree in this picture? ASSISTANT: This is the selected tree.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the person in this picture? ASSISTANT: Take a look at the person here.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[156, 156, 156,  ..., 156, 156, 156],
        [156, 156, 156,  ..., 156, 156, 156],
        [156, 156, 156,  ..., 156, 156, 156],
        ...,
        [158, 158, 158,  ..., 158, 158, 158],
        [158, 158, 158,  ..., 158, 158, 158],
        [158, 158, 158,  ..., 158, 158, 158]]), (1024, 768), ['<image>\nWhat is mountain in this image? Please give some explanation.', '<image>\nPlease highlight the snow in this image.', '<image>\nDisplay a segmentation mask for the snowboard shown in this image.', '<image>\nCould you isolate the tree in this picture?', '<image>\nCould you isolate the person in this picture?'], ['mountain', 'snow', 'snowboard', 'tree', 'person'], [{'mountain': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'snow': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True]])}, {'snowboard': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'tree': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'person': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000450414.jpg', tensor([[[-1.2959, -1.2959, -1.3130,  ...,  0.5193,  0.5193,  0.5193],
         [-1.2788, -1.2959, -1.3130,  ...,  0.5193,  0.5193,  0.5193],
         [-1.2617, -1.2788, -1.3130,  ...,  0.5193,  0.5193,  0.5193],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.1954, -1.1954, -1.2129,  ...,  0.6604,  0.6604,  0.6604],
         [-1.1779, -1.1954, -1.2129,  ...,  0.6604,  0.6604,  0.6604],
         [-1.1604, -1.1779, -1.2129,  ...,  0.6604,  0.6604,  0.6604],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.9678, -0.9678, -0.9853,  ...,  0.8797,  0.8797,  0.8797],
         [-0.9504, -0.9678, -0.9853,  ...,  0.8797,  0.8797,  0.8797],
         [-0.9330, -0.9504, -0.9853,  ...,  0.8797,  0.8797,  0.8797],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.0623, -1.0623, -1.0477,  ...,  0.5873,  0.5727,  0.5727],
         [-1.0331, -1.0185, -1.0623,  ...,  0.5873,  0.5873,  0.5727],
         [-1.1353, -1.0915, -1.1061,  ...,  0.5873,  0.5873,  0.5873],
         ...,
         [-0.3470, -0.4054, -0.4200,  ..., -0.2302, -0.2302, -0.1572],
         [-0.3616, -0.3762, -0.3470,  ..., -0.2740, -0.2302, -0.2010],
         [-0.3324, -0.3470, -0.2740,  ..., -0.2010, -0.2010, -0.2156]],

        [[-1.0017, -1.0017, -0.9867,  ...,  0.6942,  0.6792,  0.6792],
         [-0.9717, -0.9567, -1.0017,  ...,  0.6942,  0.6942,  0.6792],
         [-1.0767, -1.0317, -1.0467,  ...,  0.6942,  0.6942,  0.6942],
         ...,
         [-0.2663, -0.3264, -0.3414,  ..., -0.1463, -0.1463, -0.0712],
         [-0.2813, -0.2963, -0.2663,  ..., -0.1913, -0.1463, -0.1163],
         [-0.2513, -0.2663, -0.1913,  ..., -0.1163, -0.1163, -0.1313]],

        [[-0.7692, -0.7692, -0.7550,  ...,  0.8377,  0.8234,  0.8234],
         [-0.7408, -0.7266, -0.7692,  ...,  0.8377,  0.8377,  0.8234],
         [-0.8403, -0.7977, -0.8119,  ...,  0.8377,  0.8377,  0.8377],
         ...,
         [-0.0724, -0.1293, -0.1435,  ...,  0.0413,  0.0413,  0.1124],
         [-0.0867, -0.1009, -0.0724,  ..., -0.0013,  0.0413,  0.0698],
         [-0.0582, -0.0724, -0.0013,  ...,  0.0698,  0.0698,  0.0555]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the potential risks and safety measures to be considered in this activity? ASSISTANT: When engaging in skateboarding activities, especially while performing tricks on ramps or half pipes, there are potential risks and safety measures that should be taken into account. Some of the risks include injuries resulting from falls, collisions, or loss of control of the skateboard. Skateboarders can also suffer from sprains, fractures, and other more severe injuries if they fail to execute tricks correctly.\n\nTo minimize these risks, skateboarders should consider adopting certain safety measures, such as wearing appropriate protective gear, including helmets, wrist guards, elbow pads, and knee pads. Practicing proper techniques, warming up before skating, and gradually progressing in trick difficulty can also help prevent injuries. Furthermore, it is crucial to skateboard in a designated area like a skate park with suitable facilities and ramps designed specifically for this sport to ensure a safer environment for performing tricks and minimizing potential hazards.</s>"], tensor([], size=(0, 428, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (685, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the potential risks and safety measures to be considered in this activity? ASSISTANT: When engaging in skateboarding activities, especially while performing tricks on ramps or half pipes, there are potential risks and safety measures that should be taken into account. Some of the risks include injuries resulting from falls, collisions, or loss of control of the skateboard. Skateboarders can also suffer from sprains, fractures, and other more severe injuries if they fail to execute tricks correctly.\n\nTo minimize these risks, skateboarders should consider adopting certain safety measures, such as wearing appropriate protective gear, including helmets, wrist guards, elbow pads, and knee pads. Practicing proper techniques, warming up before skating, and gradually progressing in trick difficulty can also help prevent injuries. Furthermore, it is crucial to skateboard in a designated area like a skate park with suitable facilities and ramps designed specifically for this sport to ensure a safer environment for performing tricks and minimizing potential hazards.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat are the potential risks and safety measures to be considered in this activity? ASSISTANT: When engaging in skateboarding activities, especially while performing tricks on ramps or half pipes, there are potential risks and safety measures that should be taken into account. Some of the risks include injuries resulting from falls, collisions, or loss of control of the skateboard. Skateboarders can also suffer from sprains, fractures, and other more severe injuries if they fail to execute tricks correctly.\n\nTo minimize these risks, skateboarders should consider adopting certain safety measures, such as wearing appropriate protective gear, including helmets, wrist guards, elbow pads, and knee pads. Practicing proper techniques, warming up before skating, and gradually progressing in trick difficulty can also help prevent injuries. Furthermore, it is crucial to skateboard in a designated area like a skate park with suitable facilities and ramps designed specifically for this sport to ensure a safer environment for performing tricks and minimizing potential hazards.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000478683.jpg', tensor([[[ 0.9646,  0.3823, -0.3198,  ..., -0.2342,  0.0227,  0.1768],
         [ 0.9646,  0.3823, -0.3369,  ..., -0.2856, -0.0287,  0.1254],
         [ 0.9474,  0.3652, -0.3712,  ..., -0.3369, -0.0801,  0.0741],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.1331,  0.5903, -0.0749,  ...,  0.0651,  0.3452,  0.5203],
         [ 1.0980,  0.5553, -0.1450,  ...,  0.0126,  0.2927,  0.4678],
         [ 1.0630,  0.5028, -0.2150,  ..., -0.0399,  0.2402,  0.4153],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.6182, -0.2184, -1.2641,  ...,  0.3045,  0.5659,  0.7402],
         [ 0.6182, -0.2184, -1.2816,  ...,  0.2522,  0.5136,  0.6879],
         [ 0.6356, -0.2184, -1.2990,  ...,  0.1999,  0.4614,  0.6356],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.8063,  1.1128,  1.2442,  ..., -0.2594, -0.3470, -0.5076],
         [ 0.4559,  0.8938,  1.2296,  ..., -0.0405, -0.3178, -0.5806],
         [ 0.3391,  0.3829,  0.7771,  ..., -0.7704, -1.0477, -1.2667],
         ...,
         [ 0.2515,  0.3099,  0.3391,  ..., -1.6609, -1.6755, -1.6609],
         [ 0.0471,  0.1347, -0.0113,  ..., -1.6609, -1.6755, -1.6755],
         [-0.4784, -0.7266, -1.1207,  ..., -1.6463, -1.6463, -1.6025]],

        [[ 0.6792,  0.8893,  1.0093,  ..., -1.4519, -1.4069, -1.4519],
         [ 0.3640,  0.6642,  0.9193,  ..., -1.5420, -1.4970, -1.5420],
         [ 0.2890,  0.3640,  0.5891,  ..., -1.5720, -1.5570, -1.5570],
         ...,
         [ 0.0789,  0.0939,  0.1089,  ..., -1.6470, -1.6621, -1.6621],
         [-0.1613, -0.0862, -0.2513,  ..., -1.6470, -1.6621, -1.6621],
         [-0.6115, -0.8816, -1.2568,  ..., -1.6470, -1.6320, -1.5870]],

        [[-1.0963, -1.1816, -1.2954,  ..., -1.1247, -1.1105, -1.1674],
         [-1.1247, -1.1532, -1.2243,  ..., -1.1958, -1.1816, -1.2243],
         [-1.0963, -1.0536, -1.1247,  ..., -1.2527, -1.2527, -1.2669],
         ...,
         [ 0.5248,  0.4964,  0.4964,  ..., -1.3380, -1.3522, -1.3380],
         [ 0.2546,  0.2831,  0.1124,  ..., -1.3380, -1.3522, -1.3522],
         [-0.2289, -0.4990, -0.8830,  ..., -1.3380, -1.3238, -1.2811]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the lol left guy in this picture? ASSISTANT: Take a look at the lol left guy here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the female glasses in this image. ASSISTANT: female glasses.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the left dude in this image. ASSISTANT: I've marked the left dude for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is man on left in this image? Please elaborate your answer and explain why. ASSISTANT: Here's the area for man on left.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is lady in this image? Please give some explanation. ASSISTANT: This is the selected lady.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (691, 1024), ['<image>\nCould you isolate the lol left guy in this picture?', '<image>\nPlease highlight the female glasses in this image.', '<image>\nPlease highlight the left dude in this image.', '<image>\nWhat is man on left in this image? Please elaborate your answer and explain why.', '<image>\nWhat is lady in this image? Please give some explanation.'], ['lol left guy', 'female glasses', 'left dude', 'man on left', 'lady'], [{'lol left guy': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'female glasses': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'left dude': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'man on left': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'lady': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg')]
wandb: - 0.015 MB of 0.015 MB uploaded
>> len(batch):  6
>> batch:  >> len(batch):  6
>> batch:  [('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000324849.jpg', tensor([[[-0.2342, -0.1314, -0.0116,  ...,  0.0227, -0.0287, -0.0801],
         [-0.4054, -0.2171, -0.0116,  ..., -0.0287, -0.0458, -0.0629],
         [-0.5596, -0.3027,  0.0056,  ..., -0.0801, -0.0629, -0.0458],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.0924,  0.1352,  0.4153,  ...,  0.2577,  0.2577,  0.2577],
         [-0.2150,  0.0651,  0.4153,  ...,  0.2402,  0.2577,  0.2577],
         [-0.3375, -0.0049,  0.4328,  ...,  0.2227,  0.2577,  0.2752],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.4439,  0.6705,  0.9494,  ...,  0.9319,  0.9319,  0.9319],
         [ 0.3219,  0.6182,  0.9668,  ...,  0.9319,  0.9145,  0.9145],
         [ 0.2348,  0.5659,  1.0017,  ...,  0.9145,  0.8971,  0.8797],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[0.3683, 0.3537, 0.3537,  ..., 0.3537, 0.3537, 0.3537],
         [0.3391, 0.3391, 0.2953,  ..., 0.3099, 0.3099, 0.3099],
         [0.3245, 0.3099, 0.3245,  ..., 0.2953, 0.2807, 0.3245],
         ...,
         [0.0033, 0.0325, 0.0325,  ..., 0.2369, 0.2515, 0.2223],
         [0.0325, 0.0179, 0.0179,  ..., 0.2807, 0.2369, 0.2223],
         [0.0617, 0.0763, 0.0471,  ..., 0.2223, 0.1931, 0.2077]],

        [[0.5591, 0.5591, 0.5591,  ..., 0.4390, 0.4540, 0.4540],
         [0.5291, 0.5591, 0.5441,  ..., 0.4390, 0.4540, 0.4240],
         [0.5441, 0.4991, 0.5291,  ..., 0.4540, 0.4540, 0.4390],
         ...,
         [0.1839, 0.1989, 0.1989,  ..., 0.3340, 0.3490, 0.3490],
         [0.2439, 0.2439, 0.2289,  ..., 0.3490, 0.3490, 0.3490],
         [0.2589, 0.2740, 0.2439,  ..., 0.3190, 0.3640, 0.3340]],

        [[1.0083, 1.0225, 1.0225,  ..., 0.9514, 0.9514, 0.9372],
         [1.0083, 1.0367, 1.0225,  ..., 0.9230, 0.9372, 0.9372],
         [1.0367, 1.0083, 1.0225,  ..., 0.9372, 0.9230, 0.9088],
         ...,
         [0.7808, 0.7808, 0.8092,  ..., 0.9372, 0.9514, 0.9514],
         [0.8234, 0.8234, 0.8234,  ..., 0.9656, 0.9514, 0.9372],
         [0.8661, 0.8803, 0.8519,  ..., 0.9230, 0.9656, 0.9514]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is an important safety precaution the skier is following? ASSISTANT: An important safety precaution that the skier in the image is following is wearing a helmet. This is essential for protecting the head from potential injuries in case of falls or collisions while skiing down the slope. Considering skiing can be a risky activity, especially at high speeds or on challenging terrains, wearing a helmet significantly reduces the risk of severe head injuries and ensures a safer experience on the snow.</s>"], tensor([], size=(0, 383, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (613, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is an important safety precaution the skier is following? ASSISTANT: An important safety precaution that the skier in the image is following is wearing a helmet. This is essential for protecting the head from potential injuries in case of falls or collisions while skiing down the slope. Considering skiing can be a risky activity, especially at high speeds or on challenging terrains, wearing a helmet significantly reduces the risk of severe head injuries and ensures a safer experience on the snow.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is an important safety precaution the skier is following? ASSISTANT: An important safety precaution that the skier in the image is following is wearing a helmet. This is essential for protecting the head from potential injuries in case of falls or collisions while skiing down the slope. Considering skiing can be a risky activity, especially at high speeds or on challenging terrains, wearing a helmet significantly reduces the risk of severe head injuries and ensures a safer experience on the snow.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/r00ecQiiTySpO4peSaHRZw.jpg', tensor([[[-1.1589, -1.1760, -1.2103,  ...,  0.3138,  0.2967,  0.2796],
         [-1.1075, -1.1418, -1.1589,  ...,  0.2967,  0.2796,  0.2624],
         [-1.1075, -1.1075, -1.1247,  ...,  0.2453,  0.2453,  0.2282],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.8102, -0.7927, -0.7927,  ...,  0.8529,  0.8354,  0.8179],
         [-0.8803, -0.8803, -0.8452,  ...,  0.8880,  0.8529,  0.8354],
         [-0.8803, -0.8627, -0.8102,  ...,  0.8880,  0.8704,  0.8529],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.0256,  0.0256,  0.0256,  ...,  2.1694,  2.1346,  2.1346],
         [-0.0092, -0.0092,  0.0082,  ...,  2.1694,  2.1346,  2.1171],
         [-0.0092,  0.0082,  0.0431,  ...,  2.1520,  2.1520,  2.1171],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.0623, -1.0915, -1.1061,  ...,  0.4121,  0.4121,  0.4267],
         [-1.0623, -1.0623, -1.1061,  ...,  0.4121,  0.4267,  0.4121],
         [-1.1353, -1.0915, -0.9893,  ...,  0.4413,  0.4121,  0.3975],
         ...,
         [ 0.2369,  0.2369,  0.2369,  ..., -1.1061, -1.1207, -1.1499],
         [ 0.2661,  0.2515,  0.2369,  ..., -1.1207, -1.1353, -1.1499],
         [ 0.2369,  0.2223,  0.2515,  ..., -1.1353, -1.1499, -1.1645]],

        [[-0.8066, -0.9567, -0.9717,  ...,  0.9343,  0.9193,  0.9193],
         [-0.8967, -0.8967, -0.9117,  ...,  0.9343,  0.9343,  0.9193],
         [-0.9417, -0.8516, -0.7766,  ...,  0.9343,  0.9343,  0.9343],
         ...,
         [ 0.5891,  0.6191,  0.6341,  ..., -0.9717, -0.9717, -0.9867],
         [ 0.5891,  0.6341,  0.6341,  ..., -0.9867, -1.0017, -1.0167],
         [ 0.6041,  0.6191,  0.6491,  ..., -1.0167, -1.0167, -0.9867]],

        [[-0.1862, -0.3426, -0.3995,  ...,  1.9895,  2.0037,  2.0037],
         [-0.2715, -0.2715, -0.3426,  ...,  2.0037,  2.0179,  2.0037],
         [-0.3284, -0.2715, -0.1435,  ...,  2.0179,  2.0179,  2.0037],
         ...,
         [ 1.7193,  1.7335,  1.7193,  ..., -0.4279, -0.4279, -0.4422],
         [ 1.7051,  1.7477,  1.7477,  ..., -0.4422, -0.4279, -0.4279],
         [ 1.6909,  1.7335,  1.7620,  ..., -0.4564, -0.4706, -0.4564]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the curb in this image. ASSISTANT: Sure — this shows the curb.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the signage - store in this picture? ASSISTANT: This is the selected signage - store.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the utility pole in this image. ASSISTANT: Below you can see the utility pole.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the manhole in this image? ASSISTANT: Here's the area for manhole.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the vehicle group from the image below? ASSISTANT: You can see the vehicle group in this frame.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[64, 64, 64,  ..., 61, 61, 61],
        [64, 64, 64,  ..., 61, 61, 61],
        [64, 64, 64,  ..., 61, 61, 61],
        ...,
        [62, 62, 62,  ..., 21, 21, 21],
        [62, 62, 62,  ..., 21, 21, 21],
        [62, 62, 62,  ..., 21, 21, 21]]), (768, 1024), ['<image>\nPlease segment the curb in this image.', '<image>\nCould you isolate the signage - store in this picture?', '<image>\nPlease segment the utility pole in this image.', '<image>\nCan you segment the manhole in this image?', '<image>\nWould you please extract the vehicle group from the image below?'], ['curb', 'signage - store', 'utility pole', 'manhole', 'vehicle group'], [{'curb': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'signage - store': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'utility pole': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'manhole': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'vehicle group': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/refer_seg/images/mscoco/images/train2014/COCO_train2014_000000155543.jpg', tensor([[[-1.1589, -1.1589, -1.1589,  ..., -1.2445, -1.2617, -1.2788],
         [-1.1589, -1.1589, -1.1589,  ..., -1.2445, -1.2617, -1.2788],
         [-1.1589, -1.1589, -1.1589,  ..., -1.2445, -1.2617, -1.2617],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-0.1450, -0.1450, -0.1450,  ..., -0.3200, -0.3375, -0.3550],
         [-0.1450, -0.1450, -0.1450,  ..., -0.3200, -0.3375, -0.3550],
         [-0.1450, -0.1450, -0.1450,  ..., -0.3200, -0.3375, -0.3375],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 1.3851,  1.3851,  1.3851,  ...,  1.0888,  1.0714,  1.0539],
         [ 1.3851,  1.3851,  1.3851,  ...,  1.0888,  1.0714,  1.0539],
         [ 1.3851,  1.3851,  1.3851,  ...,  1.0888,  1.0714,  1.0714],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.9018, -0.9018, -0.8726,  ..., -0.9310, -0.9310, -0.9602],
         [-0.9018, -0.9018, -0.8580,  ..., -0.9456, -0.9456, -0.9456],
         [-0.8872, -0.8726, -0.8288,  ..., -0.9018, -0.9164, -0.9310],
         ...,
         [ 0.6749,  0.6895,  0.7479,  ..., -0.5222,  0.1931,  0.6895],
         [ 0.7333,  0.7479,  0.7917,  ...,  0.1639, -0.3908,  0.5727],
         [ 0.7187,  0.7187,  0.7479,  ...,  0.4413, -0.4930, -0.1864]],

        [[-0.0562, -0.0562, -0.0562,  ..., -0.1463, -0.1613, -0.1613],
         [-0.0562, -0.0562, -0.0412,  ..., -0.1463, -0.1463, -0.1463],
         [-0.0412, -0.0412, -0.0262,  ..., -0.1313, -0.1463, -0.1313],
         ...,
         [ 0.7992,  0.7992,  0.8442,  ..., -0.4314,  0.4090,  0.8142],
         [ 0.8442,  0.8442,  0.8893,  ...,  0.2139, -0.2813,  0.7092],
         [ 0.8142,  0.8142,  0.8442,  ...,  0.4240, -0.3864,  0.0038]],

        [[ 1.1647,  1.1647,  1.1647,  ...,  1.0225,  1.0225,  1.0083],
         [ 1.1647,  1.1647,  1.1789,  ...,  1.0225,  1.0225,  1.0225],
         [ 1.1789,  1.1789,  1.2074,  ...,  1.0510,  1.0367,  1.0367],
         ...,
         [ 0.8945,  0.8945,  0.9372,  ..., -0.1151,  0.5532,  0.8377],
         [ 0.9372,  0.9372,  0.9656,  ...,  0.4253,  0.0129,  0.7950],
         [ 0.8945,  0.8945,  0.9088,  ...,  0.6812, -0.0013,  0.2404]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the white shirt from the image below? ASSISTANT: white shirt.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the red shirt in this image. ASSISTANT: The red shirt is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the guy in white from the provided image. ASSISTANT: This view focuses on the guy in white.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the man in red shirt shown in this image. ASSISTANT: I've marked the man in red shirt for you.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the skateboarder in the air in this picture? ASSISTANT: You can see the skateboarder in the air in this frame.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (752, 1024), ['<image>\nWould you please extract the white shirt from the image below?', '<image>\nPlease highlight the red shirt in this image.', '<image>\nSegment the guy in white from the provided image.', '<image>\nDisplay a segmentation mask for the man in red shirt shown in this image.', '<image>\nCould you identify and segment out the skateboarder in the air in this picture?'], ['white shirt', 'red shirt', 'guy in white', 'man in red shirt', 'skateboarder in the air'], [{'white shirt': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'red shirt': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'guy in white': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'man in red shirt': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'skateboarder in the air': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], [], False, 'refer_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/mapillary/training/images/JVHC1CtC9jh0Bqrmlfp0PA.jpg', tensor([[[-2.0494, -2.0665, -2.0152,  ...,  0.4679,  0.4679,  0.4851],
         [-2.0494, -2.0665, -2.0494,  ...,  0.4679,  0.4679,  0.4508],
         [-2.0323, -2.0494, -2.0494,  ...,  0.4679,  0.4508,  0.4337],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.8782, -1.8957, -1.8431,  ...,  1.5182,  1.5182,  1.5357],
         [-1.8782, -1.8957, -1.8782,  ...,  1.5182,  1.5182,  1.5007],
         [-1.8606, -1.8782, -1.8782,  ...,  1.5182,  1.5007,  1.4832],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.7173, -1.7347, -1.6824,  ...,  2.3786,  2.3786,  2.3960],
         [-1.7173, -1.7347, -1.7173,  ...,  2.3786,  2.3786,  2.3611],
         [-1.6999, -1.7173, -1.7173,  ...,  2.3786,  2.3611,  2.3437],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-1.7485, -1.7485, -1.7777,  ...,  0.3829,  0.4559,  0.4705],
         [-1.6463, -1.7339, -1.7777,  ...,  0.4705,  0.4997,  0.4997],
         [-1.5733, -1.6901, -1.7339,  ...,  0.5143,  0.5435,  0.5435],
         ...,
         [-1.6025, -1.5295, -1.4857,  ..., -0.8434, -0.8288, -0.8434],
         [-1.4857, -1.4419, -1.3981,  ..., -0.8288, -0.8434, -0.8434],
         [-1.4857, -1.4711, -1.4419,  ..., -0.8288, -0.8288, -0.8288]],

        [[-1.6320, -1.6771, -1.7071,  ...,  1.2344,  1.3095,  1.3245],
         [-1.5120, -1.6170, -1.6771,  ...,  1.3245,  1.3395,  1.3545],
         [-1.4219, -1.5420, -1.6020,  ...,  1.3545,  1.3695,  1.3695],
         ...,
         [-1.4669, -1.4069, -1.3619,  ..., -0.6865, -0.6715, -0.6865],
         [-1.3619, -1.3169, -1.3019,  ..., -0.6715, -0.6865, -0.7016],
         [-1.4069, -1.3919, -1.3619,  ..., -0.7016, -0.7016, -0.7016]],

        [[-1.4518, -1.4518, -1.4660,  ...,  1.8473,  1.9184,  1.9326],
         [-1.3949, -1.4518, -1.4518,  ...,  1.9184,  1.9184,  1.9326],
         [-1.3522, -1.4518, -1.4376,  ...,  1.9468,  1.9468,  1.9468],
         ...,
         [-1.2385, -1.1958, -1.1816,  ..., -0.4706, -0.4564, -0.4564],
         [-1.1674, -1.1247, -1.1105,  ..., -0.4564, -0.4706, -0.4422],
         [-1.1674, -1.1816, -1.1532,  ..., -0.4706, -0.4706, -0.4422]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the curb in this image. ASSISTANT: The curb is presented here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the fence shown in this image. ASSISTANT: The fence portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is traffic island in this image? Please elaborate your answer and explain why. ASSISTANT: Sure — this shows the traffic island.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is road in this image? Please elaborate your answer and explain why. ASSISTANT: Sure — this shows the road.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the lane marking - straight line from the provided image. ASSISTANT: The lane marking - straight line portion is displayed below.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ...,  True,  True,  True],
         [False, False, False,  ...,  True,  True,  True],
         [False, False, False,  ...,  True,  True,  True]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[64, 64, 64,  ..., 61, 61, 61],
        [64, 64, 64,  ..., 61, 61, 61],
        [64, 64, 64,  ..., 61, 61, 61],
        ...,
        [ 4,  4,  4,  ..., 21, 21, 21],
        [ 4,  4,  4,  ..., 21, 21, 21],
        [ 4,  4,  4,  ..., 21, 21, 21]]), (768, 1024), ['<image>\nPlease highlight the curb in this image.', '<image>\nDisplay a segmentation mask for the fence shown in this image.', '<image>\nWhat is traffic island in this image? Please elaborate your answer and explain why.', '<image>\nWhat is road in this image? Please elaborate your answer and explain why.', '<image>\nSegment the lane marking - straight line from the provided image.'], ['curb', 'fence', 'traffic island', 'road', 'lane marking - straight line'], [{'curb': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False]])}, {'fence': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'traffic island': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'road': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ...,  True,  True,  True],
        [False, False, False,  ...,  True,  True,  True],
        [False, False, False,  ...,  True,  True,  True]])}, {'lane marking - straight line': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000333247.jpg', tensor([[[0.8104, 0.8104, 0.8276,  ..., 0.0227, 0.0227, 0.0227],
         [0.8104, 0.8104, 0.8276,  ..., 0.0398, 0.0398, 0.0227],
         [0.8104, 0.8104, 0.8276,  ..., 0.0741, 0.0569, 0.0398],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.9580, 0.9580, 0.9755,  ..., 0.1527, 0.1527, 0.1527],
         [0.9580, 0.9580, 0.9755,  ..., 0.1702, 0.1702, 0.1527],
         [0.9580, 0.9580, 0.9755,  ..., 0.2052, 0.1877, 0.1702],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.1411, 1.1411, 1.1585,  ..., 0.3742, 0.3742, 0.3742],
         [1.1411, 1.1411, 1.1585,  ..., 0.3916, 0.3916, 0.3742],
         [1.1411, 1.1411, 1.1585,  ..., 0.4265, 0.4091, 0.3916],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.3318,  1.3318,  1.3172,  ...,  0.6603,  0.6603,  0.6603],
         [ 1.3318,  1.3464,  1.3172,  ...,  0.6311,  0.6165,  0.6019],
         [ 1.3464,  1.3464,  1.3318,  ...,  0.6165,  0.6165,  0.6019],
         ...,
         [ 0.6311,  0.6749,  0.6895,  ..., -1.5733, -1.5003, -1.5587],
         [ 0.6457,  0.6457,  0.6457,  ..., -1.5879, -1.5587, -1.5003],
         [ 0.6311,  0.6603,  0.6457,  ..., -1.2667, -1.1061, -1.0331]],

        [[ 1.4596,  1.4596,  1.4446,  ...,  0.7392,  0.7392,  0.7392],
         [ 1.4596,  1.4746,  1.4446,  ...,  0.7092,  0.6942,  0.6942],
         [ 1.4746,  1.4746,  1.4596,  ...,  0.7242,  0.7242,  0.6942],
         ...,
         [ 0.7842,  0.7842,  0.7692,  ..., -1.5570, -1.5120, -1.5570],
         [ 0.7692,  0.7692,  0.7692,  ..., -1.5570, -1.5570, -1.4970],
         [ 0.6041,  0.7542,  0.7692,  ..., -1.3469, -1.2268, -1.1518]],

        [[ 1.5629,  1.5629,  1.5487,  ...,  0.8945,  0.8945,  0.8803],
         [ 1.5771,  1.5771,  1.5487,  ...,  0.8661,  0.8519,  0.8377],
         [ 1.5771,  1.5771,  1.5629,  ...,  0.8661,  0.8661,  0.8519],
         ...,
         [ 0.9514,  0.9230,  0.9230,  ..., -1.3807, -1.3096, -1.3380],
         [ 0.9372,  0.9372,  0.9088,  ..., -1.3238, -1.3380, -1.3380],
         [ 0.9088,  0.9230,  0.8945,  ..., -1.1958, -1.1247, -1.1247]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat advice would you give to someone organizing this space? ASSISTANT: To effectively organize this cluttered space filled with boxes, furniture, and various items, I would suggest the following steps:\n\n1. Categorize the items: Sort through the boxes and create categories for the contents, such as books, electronics, kitchenware, and clothing. This will help you determine where the items should be placed within the house and make the unpacking process more efficient.\n\n2. Prioritize high-use items: Identify and prioritize items that will be needed immediately or frequently. For instance, the laptop and essential kitchenware should be unpacked and placed in easily accessible locations.\n\n3. Start with larger items: Begin by arranging larger pieces of furniture, such as the chair and any other bulky items. This will give you a better idea of the available space for other items.\n\n4. Find designated spaces: Allocate specific areas or storage solutions for items like books and the potted plant. Consider adding shelves for the books and finding a suitable spot with sufficient light for the potted plant.\n\n5. Dispose of or donate unneeded items: As you unpack and organize, identify any items that are no longer needed or can be donated. This will help declutter the space further and create a more organized environment.\n\n6. Maintain organization: Once you have finished unpacking and organizing, establish a routine to maintain the order within the space to prevent future clutter.\n\nBy following these steps, the cluttered room can be transformed into a more organized and functional space that is easier to navigate and maintain.</s>"], tensor([], size=(0, 480, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat advice would you give to someone organizing this space? ASSISTANT: To effectively organize this cluttered space filled with boxes, furniture, and various items, I would suggest the following steps:\n\n1. Categorize the items: Sort through the boxes and create categories for the contents, such as books, electronics, kitchenware, and clothing. This will help you determine where the items should be placed within the house and make the unpacking process more efficient.\n\n2. Prioritize high-use items: Identify and prioritize items that will be needed immediately or frequently. For instance, the laptop and essential kitchenware should be unpacked and placed in easily accessible locations.\n\n3. Start with larger items: Begin by arranging larger pieces of furniture, such as the chair and any other bulky items. This will give you a better idea of the available space for other items.\n\n4. Find designated spaces: Allocate specific areas or storage solutions for items like books and the potted plant. Consider adding shelves for the books and finding a suitable spot with sufficient light for the potted plant.\n\n5. Dispose of or donate unneeded items: As you unpack and organize, identify any items that are no longer needed or can be donated. This will help declutter the space further and create a more organized environment.\n\n6. Maintain organization: Once you have finished unpacking and organizing, establish a routine to maintain the order within the space to prevent future clutter.\n\nBy following these steps, the cluttered room can be transformed into a more organized and functional space that is easier to navigate and maintain.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat advice would you give to someone organizing this space? ASSISTANT: To effectively organize this cluttered space filled with boxes, furniture, and various items, I would suggest the following steps:\n\n1. Categorize the items: Sort through the boxes and create categories for the contents, such as books, electronics, kitchenware, and clothing. This will help you determine where the items should be placed within the house and make the unpacking process more efficient.\n\n2. Prioritize high-use items: Identify and prioritize items that will be needed immediately or frequently. For instance, the laptop and essential kitchenware should be unpacked and placed in easily accessible locations.\n\n3. Start with larger items: Begin by arranging larger pieces of furniture, such as the chair and any other bulky items. This will give you a better idea of the available space for other items.\n\n4. Find designated spaces: Allocate specific areas or storage solutions for items like books and the potted plant. Consider adding shelves for the books and finding a suitable spot with sufficient light for the potted plant.\n\n5. Dispose of or donate unneeded items: As you unpack and organize, identify any items that are no longer needed or can be donated. This will help declutter the space further and create a more organized environment.\n\n6. Maintain organization: Once you have finished unpacking and organizing, establish a routine to maintain the order within the space to prevent future clutter.\n\nBy following these steps, the cluttered room can be transformed into a more organized and functional space that is easier to navigate and maintain.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/vlpart/pascal_part/VOCdevkit/VOC2010/JPEGImages/2008_002357.jpg', tensor([[[0.1939, 0.2282, 0.2967,  ..., 0.7933, 0.7762, 0.7591],
         [0.1939, 0.2282, 0.2967,  ..., 0.8104, 0.7933, 0.7762],
         [0.2111, 0.2282, 0.2796,  ..., 0.8276, 0.8447, 0.8276],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.3102, 0.3452, 0.4153,  ..., 0.8529, 0.8529, 0.8529],
         [0.3102, 0.3452, 0.3978,  ..., 0.8704, 0.8704, 0.8704],
         [0.3277, 0.3277, 0.3627,  ..., 0.9230, 0.9230, 0.9230],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[0.4439, 0.4788, 0.5485,  ..., 1.2457, 1.2282, 1.2282],
         [0.4614, 0.4962, 0.5485,  ..., 1.2631, 1.2631, 1.2631],
         [0.4962, 0.5136, 0.5485,  ..., 1.3154, 1.3328, 1.3328],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[1.2150, 1.1420, 1.0836,  ..., 1.0690, 1.1128, 1.1128],
         [1.2880, 1.2150, 1.1420,  ..., 1.0836, 1.1128, 1.1274],
         [1.3026, 1.2588, 1.2004,  ..., 1.1420, 1.1420, 1.1566],
         ...,
         [1.1420, 1.1420, 1.1274,  ..., 0.7625, 0.6457, 0.7333],
         [1.1566, 1.1274, 1.1128,  ..., 0.6311, 0.5289, 0.5435],
         [1.0836, 1.0836, 1.0544,  ..., 0.6311, 0.6457, 0.4997]],

        [[1.2945, 1.2044, 1.1294,  ..., 1.1894, 1.2645, 1.2645],
         [1.3995, 1.3095, 1.2194,  ..., 1.2044, 1.2645, 1.2795],
         [1.4295, 1.3845, 1.3245,  ..., 1.2645, 1.2795, 1.2795],
         ...,
         [1.3245, 1.3245, 1.3095,  ..., 0.8893, 0.7842, 0.8893],
         [1.3245, 1.2945, 1.2795,  ..., 0.7392, 0.6491, 0.6642],
         [1.2344, 1.2194, 1.2344,  ..., 0.7542, 0.7692, 0.6191]],

        [[1.3638, 1.2785, 1.2074,  ..., 1.4776, 1.5487, 1.5629],
         [1.4918, 1.3922, 1.3211,  ..., 1.4918, 1.5487, 1.5771],
         [1.5344, 1.4633, 1.4065,  ..., 1.5487, 1.5487, 1.5629],
         ...,
         [1.5629, 1.5487, 1.5487,  ..., 1.1363, 1.0367, 1.1221],
         [1.5629, 1.5344, 1.5344,  ..., 0.9656, 0.8661, 0.8945],
         [1.4776, 1.4633, 1.4918,  ..., 0.9656, 0.9799, 0.8377]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the the arm of the person in this image? ASSISTANT: This is the selected the arm of the person.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the person arm from the image below? ASSISTANT: person arm.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the person torso in this image. ASSISTANT: This is the selected person torso.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the the mouth of the person from the image below? ASSISTANT: Here's where the the mouth of the person appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the the eye of the person in this image. ASSISTANT: The region corresponding to the the eye of the person is shown.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['<image>\nCan you segment the the arm of the person in this image?', '<image>\nWould you please extract the person arm from the image below?', '<image>\nPlease segment the person torso in this image.', '<image>\nWould you please extract the the mouth of the person from the image below?', '<image>\nPlease segment the the eye of the person in this image.'], ['the arm of the person', 'person arm', 'person torso', 'the mouth of the person', 'the eye of the person'], [{'the arm of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person arm': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'person torso': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the mouth of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the eye of the person': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg')]
[('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000335459.jpg', tensor([[[-0.1828, -0.1828, -0.1828,  ..., -0.5082, -0.5082, -0.5082],
         [-0.1486, -0.1657, -0.1657,  ..., -0.5424, -0.5424, -0.5424],
         [-0.1143, -0.1314, -0.1486,  ..., -0.5767, -0.5938, -0.5938],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.5903,  0.5903,  0.5728,  ...,  0.1702,  0.1352,  0.1176],
         [ 0.5903,  0.5903,  0.5728,  ...,  0.1702,  0.1352,  0.1176],
         [ 0.6078,  0.5903,  0.5728,  ...,  0.1877,  0.1527,  0.1176],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[ 0.9668,  0.9668,  0.9494,  ...,  0.6182,  0.6008,  0.5834],
         [ 0.9668,  0.9668,  0.9494,  ...,  0.6008,  0.5659,  0.5485],
         [ 0.9842,  0.9668,  0.9494,  ...,  0.5659,  0.5311,  0.5136],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[ 0.2515,  0.2515,  0.2369,  ..., -0.5514, -0.5368, -0.5222],
         [ 0.2223,  0.2223,  0.2077,  ..., -0.5514, -0.5368, -0.5222],
         [ 0.1931,  0.1931,  0.1931,  ..., -0.5368, -0.5368, -0.5222],
         ...,
         [-0.9456, -0.8726, -0.8434,  ..., -0.9602, -0.8872, -0.8872],
         [-0.9018, -0.8872, -0.8726,  ..., -1.0769, -1.0477, -0.9893],
         [-0.9018, -0.9310, -0.9310,  ..., -1.0331, -1.0769, -1.1061]],

        [[ 0.8442,  0.8442,  0.8442,  ...,  0.2439,  0.2439,  0.2439],
         [ 0.8442,  0.8593,  0.8743,  ...,  0.2439,  0.2589,  0.2589],
         [ 0.8743,  0.8593,  0.8743,  ...,  0.2589,  0.2589,  0.2740],
         ...,
         [-1.0317, -0.9717, -0.9867,  ..., -1.1218, -1.0767, -1.0918],
         [-0.9867, -0.9867, -1.0017,  ..., -1.2118, -1.2118, -1.1668],
         [-1.0017, -1.0167, -1.0317,  ..., -1.1668, -1.2418, -1.2718]],

        [[ 0.9372,  0.9941,  1.0225,  ...,  0.5959,  0.6244,  0.6101],
         [ 0.9230,  0.9230,  0.9514,  ...,  0.5532,  0.5675,  0.5817],
         [ 0.9514,  0.9514,  0.9656,  ...,  0.5959,  0.5959,  0.5959],
         ...,
         [-0.8261, -0.7834, -0.8119,  ..., -0.9967, -0.9825, -1.0252],
         [-0.7266, -0.7408, -0.7977,  ..., -1.0821, -1.0963, -1.0821],
         [-0.7123, -0.7692, -0.7834,  ..., -1.0394, -1.0963, -1.1247]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCan you segment the refrigerator in this image? ASSISTANT: This is the selected refrigerator.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the book in this picture? ASSISTANT: Below you can see the book.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nSegment the cabinet from the provided image. ASSISTANT: This view focuses on the cabinet.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the cardboard in this image. ASSISTANT: I've marked the cardboard for you.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255],
        ...,
        [255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ..., 255, 255, 255]]), (768, 1024), ['<image>\nCan you segment the refrigerator in this image?', '<image>\nCould you identify and segment out the book in this picture?', '<image>\nSegment the cabinet from the provided image.', '<image>\nPlease highlight the cardboard in this image.'], ['refrigerator', 'book', 'cabinet', 'cardboard'], [{'refrigerator': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'book': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'cabinet': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'cardboard': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000368370.jpg', tensor([[[0.7933, 0.6906, 0.5707,  ..., 0.6221, 0.7591, 0.8789],
         [0.7248, 0.6221, 0.5022,  ..., 0.5022, 0.6392, 0.7591],
         [0.6392, 0.5364, 0.3994,  ..., 0.3309, 0.4679, 0.5878],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.9559, 1.8508, 1.7283,  ..., 1.7283, 1.8333, 1.9209],
         [1.8683, 1.7633, 1.6408,  ..., 1.6057, 1.7108, 1.8158],
         [1.7633, 1.6583, 1.5182,  ..., 1.4307, 1.5532, 1.6758],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.5877, 2.4831, 2.3611,  ..., 2.3786, 2.5006, 2.6051],
         [2.5180, 2.4134, 2.2740,  ..., 2.2566, 2.3960, 2.5006],
         [2.4134, 2.3088, 2.1694,  ..., 2.1171, 2.2566, 2.3786],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 0.5581,  0.5581,  0.5727,  ...,  0.5143,  0.5289,  0.5289],
         [ 0.5143,  0.4851,  0.5143,  ...,  0.4851,  0.4851,  0.4851],
         [ 0.4997,  0.5143,  0.5289,  ...,  0.4997,  0.4851,  0.4997],
         ...,
         [-1.6025, -1.6025, -1.6025,  ..., -1.7339, -1.7631, -1.7631],
         [-1.6463, -1.6755, -1.6317,  ..., -1.7047, -1.7193, -1.7485],
         [-1.7339, -1.7485, -1.6901,  ..., -1.6025, -1.6463, -1.7193]],

        [[ 1.4746,  1.4746,  1.4746,  ...,  1.4746,  1.4896,  1.4896],
         [ 1.4145,  1.3845,  1.4145,  ...,  1.4145,  1.4145,  1.4145],
         [ 1.4145,  1.4295,  1.4446,  ...,  1.4145,  1.3995,  1.4145],
         ...,
         [-1.3619, -1.3619, -1.3769,  ..., -1.6921, -1.7221, -1.7221],
         [-1.4519, -1.4519, -1.4369,  ..., -1.6621, -1.6771, -1.7071],
         [-1.6020, -1.5870, -1.5270,  ..., -1.5570, -1.6020, -1.6771]],

        [[ 1.9610,  1.9753,  2.0037,  ...,  1.9326,  1.9610,  1.9753],
         [ 1.9610,  1.9468,  1.9753,  ...,  1.8899,  1.9042,  1.8899],
         [ 1.8757,  1.9184,  1.9468,  ...,  1.9184,  1.9042,  1.8899],
         ...,
         [-1.0963, -1.1105, -1.1389,  ..., -1.4233, -1.4518, -1.4518],
         [-1.1674, -1.1816, -1.1816,  ..., -1.3949, -1.4091, -1.4376],
         [-1.3238, -1.3380, -1.2527,  ..., -1.3380, -1.3665, -1.4233]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the man wearing while holding the surfboard? ASSISTANT: The man is wearing a wet suit while holding the surfboard.</s>USER: What is the man doing on the beach? ASSISTANT: The man is standing and walking near the ocean, carrying his white surfboard.</s>USER: What is the location of the man relative to the water? ASSISTANT: The man is near the water or near the ocean, carrying his surfboard on the beach.</s>USER: What color is the man's surfboard? ASSISTANT: The man's surfboard is white.</s>USER: What might be the man's intention with the surfboard? ASSISTANT: Given that the man is wearing a wet suit and holding a surfboard on the beach, it is likely that his intention is to go surfing. Surfing is a popular water sport where a person rides on the face of a moving wave with the help of a surfboard, and the wet suit is designed to provide thermal insulation and protection from the cold water temperatures, abrasions, and other elements. \n\nThe combination of the man's attire and his presence on the beach with a surfboard suggests that he is either preparing to get into the water to catch some waves or has recently finished surfing and is now returning to the shore. Surfing requires a suitable combination of wave size, water conditions, and weather, so it's possible that the man has been waiting for the right moment to venture into the ocean or has enjoyed a good session in the waves and is now taking a break. Regardless of his specific plan or timing, it is evident from the image that the man's primary focus and interest are related to surfing.</s>"], tensor([], size=(0, 429, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (686, 1024), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the man wearing while holding the surfboard? ASSISTANT: The man is wearing a wet suit while holding the surfboard.</s>USER: What is the man doing on the beach? ASSISTANT: The man is standing and walking near the ocean, carrying his white surfboard.</s>USER: What is the location of the man relative to the water? ASSISTANT: The man is near the water or near the ocean, carrying his surfboard on the beach.</s>USER: What color is the man's surfboard? ASSISTANT: The man's surfboard is white.</s>USER: What might be the man's intention with the surfboard? ASSISTANT: Given that the man is wearing a wet suit and holding a surfboard on the beach, it is likely that his intention is to go surfing. Surfing is a popular water sport where a person rides on the face of a moving wave with the help of a surfboard, and the wet suit is designed to provide thermal insulation and protection from the cold water temperatures, abrasions, and other elements. \n\nThe combination of the man's attire and his presence on the beach with a surfboard suggests that he is either preparing to get into the water to catch some waves or has recently finished surfing and is now returning to the shore. Surfing requires a suitable combination of wave size, water conditions, and weather, so it's possible that the man has been waiting for the right moment to venture into the ocean or has enjoyed a good session in the waves and is now taking a break. Regardless of his specific plan or timing, it is evident from the image that the man's primary focus and interest are related to surfing.</s>"], ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is the man wearing while holding the surfboard? ASSISTANT: The man is wearing a wet suit while holding the surfboard.</s>USER: What is the man doing on the beach? ASSISTANT: The man is standing and walking near the ocean, carrying his white surfboard.</s>USER: What is the location of the man relative to the water? ASSISTANT: The man is near the water or near the ocean, carrying his surfboard on the beach.</s>USER: What color is the man's surfboard? ASSISTANT: The man's surfboard is white.</s>USER: What might be the man's intention with the surfboard? ASSISTANT: Given that the man is wearing a wet suit and holding a surfboard on the beach, it is likely that his intention is to go surfing. Surfing is a popular water sport where a person rides on the face of a moving wave with the help of a surfboard, and the wet suit is designed to provide thermal insulation and protection from the cold water temperatures, abrasions, and other elements. \n\nThe combination of the man's attire and his presence on the beach with a surfboard suggests that he is either preparing to get into the water to catch some waves or has recently finished surfing and is now returning to the shore. Surfing requires a suitable combination of wave size, water conditions, and weather, so it's possible that the man has been waiting for the right moment to venture into the ocean or has enjoyed a good session in the waves and is now taking a break. Regardless of his specific plan or timing, it is evident from the image that the man's primary focus and interest are related to surfing.</s>"], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000272735.jpg', tensor([[[-1.4500, -1.2103, -0.9192,  ..., -1.4158, -1.4672, -1.4843],
         [-1.5699, -1.3987, -1.2103,  ..., -1.4843, -1.5185, -1.5185],
         [-1.6384, -1.5528, -1.4843,  ..., -1.5528, -1.5528, -1.5528],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.5280, -1.2654, -0.9153,  ..., -1.3004, -1.3354, -1.3354],
         [-1.5455, -1.3880, -1.1604,  ..., -1.3529, -1.3704, -1.3704],
         [-1.5105, -1.4580, -1.4055,  ..., -1.3880, -1.4055, -1.4055],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.1073, -0.8633, -0.5844,  ..., -1.1944, -1.3687, -1.4733],
         [-1.1247, -0.9678, -0.7936,  ..., -1.3687, -1.4559, -1.5081],
         [-1.1073, -1.0550, -0.9853,  ..., -1.5430, -1.5430, -1.5256],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[-0.7850, -0.6244, -0.7558,  ..., -1.0769, -0.9602, -0.8142],
         [-1.1207, -0.8580, -0.7996,  ..., -1.1645, -0.9602, -0.9602],
         [-1.0769, -1.2083, -1.2959,  ..., -1.0915, -1.0769, -0.9748],
         ...,
         [ 0.6311,  0.6311,  0.6165,  ...,  0.7333,  0.7041,  0.7187],
         [ 0.6311,  0.6165,  0.6311,  ...,  0.7041,  0.6895,  0.7187],
         [ 0.6457,  0.6165,  0.6165,  ...,  0.6749,  0.6749,  0.6895]],

        [[-0.5815, -0.4464, -0.7766,  ..., -1.0767, -0.9117, -0.7466],
         [-0.9567, -0.6865, -0.6565,  ..., -1.1518, -0.8967, -0.8967],
         [-0.9417, -0.9717, -1.0167,  ..., -1.0767, -1.0167, -0.8816],
         ...,
         [ 0.8142,  0.8142,  0.7992,  ...,  0.8593,  0.8442,  0.8593],
         [ 0.7992,  0.8292,  0.8292,  ...,  0.8743,  0.8743,  0.8743],
         [ 0.8292,  0.8442,  0.8292,  ...,  0.8893,  0.9043,  0.8743]],

        [[-0.6981, -0.7834, -0.5133,  ..., -0.9114, -0.9967, -0.8830],
         [-1.0110, -0.7123, -0.6412,  ..., -0.9683, -0.7977, -0.7834],
         [-0.9683, -0.9967, -1.2669,  ..., -0.8261, -0.8403, -0.8545],
         ...,
         [ 0.8519,  0.8519,  0.8377,  ...,  0.9514,  0.9088,  0.9230],
         [ 0.8377,  0.8519,  0.8661,  ...,  0.9230,  0.9230,  0.9372],
         [ 0.8519,  0.8519,  0.8519,  ...,  0.9230,  0.9230,  0.9230]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWould you please extract the handbag from the image below? ASSISTANT: Below you can see the handbag.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the the handle of the handbag in this picture. ASSISTANT: Sure — this shows the the handle of the handbag.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the handbag body in this picture? ASSISTANT: The region corresponding to the handbag body is shown.</s>"], tensor([[[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]],

        [[0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         ...,
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0],
         [0, 0, 0,  ..., 0, 0, 0]]], dtype=torch.uint8), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (683, 1024), ['<image>\nWould you please extract the handbag from the image below?', '<image>\nPlease point out the the handle of the handbag in this picture.', '<image>\nCould you identify and segment out the handbag body in this picture?'], ['handbag', 'the handle of the handbag', 'handbag body'], [{'handbag': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'the handle of the handbag': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}, {'handbag body': tensor([[0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        ...,
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0],
        [0, 0, 0,  ..., 0, 0, 0]], dtype=torch.uint8)}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000249309.jpg', tensor([[[1.5810, 1.5810, 1.5982,  ..., 1.7694, 1.7523, 1.7352],
         [1.5982, 1.5982, 1.5982,  ..., 1.7694, 1.7523, 1.7352],
         [1.6153, 1.6153, 1.6153,  ..., 1.7694, 1.7523, 1.7352],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[1.9384, 1.9384, 1.9559,  ..., 2.2710, 2.2360, 2.2185],
         [1.9559, 1.9559, 1.9559,  ..., 2.2710, 2.2360, 2.2185],
         [1.9734, 1.9734, 1.9734,  ..., 2.2535, 2.2360, 2.2185],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]],

        [[2.5006, 2.5006, 2.5180,  ..., 2.6051, 2.6051, 2.6051],
         [2.5180, 2.5180, 2.5180,  ..., 2.6226, 2.6051, 2.6051],
         [2.5354, 2.5354, 2.5354,  ..., 2.6400, 2.6226, 2.6051],
         ...,
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000],
         [0.0000, 0.0000, 0.0000,  ..., 0.0000, 0.0000, 0.0000]]]), tensor([[[ 1.7114,  1.7260,  1.7260,  ...,  1.9303,  1.9303,  1.9303],
         [ 1.7114,  1.7114,  1.7260,  ...,  1.9303,  1.9303,  1.9303],
         [ 1.7114,  1.7114,  1.7114,  ...,  1.9303,  1.9303,  1.9303],
         ...,
         [-0.9310, -0.9310, -0.8580,  ..., -0.7996, -0.9018, -0.5806],
         [-0.8872, -0.9018, -0.9018,  ..., -0.7704, -0.6244, -0.4200],
         [-0.8872, -0.9310, -0.9310,  ..., -0.9893, -0.6536, -0.6244]],

        [[ 1.8798,  1.8948,  1.9098,  ...,  2.0749,  2.0749,  2.0749],
         [ 1.8798,  1.8798,  1.8948,  ...,  2.0749,  2.0749,  2.0749],
         [ 1.8798,  1.8798,  1.8798,  ...,  2.0749,  2.0749,  2.0749],
         ...,
         [-0.7316, -0.7316, -0.7016,  ..., -0.7466, -0.8066, -0.4164],
         [-0.7016, -0.7466, -0.7466,  ..., -0.6715, -0.4914, -0.3714],
         [-0.7466, -0.7766, -0.7766,  ..., -0.8967, -0.5665, -0.5965]],

        [[ 2.1317,  2.1459,  2.1317,  ...,  2.1459,  2.1459,  2.1459],
         [ 2.1317,  2.1317,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         [ 2.1317,  2.1317,  2.1459,  ...,  2.1459,  2.1459,  2.1459],
         ...,
         [-0.1862, -0.2715, -0.1862,  ..., -0.3142, -0.6697, -0.3000],
         [-0.2004, -0.1720, -0.1293,  ..., -0.4706, -0.4279,  0.0271],
         [-0.2289, -0.2715, -0.1578,  ..., -0.7692, -0.2146, -0.1151]]]), ['A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human\'s questions. USER: <image>\nWhat color is the traffic light in the image? ASSISTANT: The traffic light in the image is green.</s>USER: Where is the traffic light located in the scene? ASSISTANT: The traffic light is located at a wide intersection in the crosswalk area, within a street intersection that also features a crosswalk and a street sign.</s>USER: Can you describe the surroundings of the traffic light? ASSISTANT: The surroundings of the traffic light include a city street, a tall house sitting between two roads, and various traffic lights and signs. In the background, there are trees that add a touch of greenery to the urban environment.</s>USER: What does the presence of the green traffic light indicate for pedestrians and drivers? ASSISTANT: The green traffic light indicates that it is currently safe for vehicles to pass through the intersection as they have the right-of-way. Pedestrians should be cautious while crossing the street during this time as the green light permits vehicles to move forward. However, this doesn\'t necessarily mean that pedestrians aren\'t allowed to cross, but they should be more vigilant and watch for incoming vehicles, especially in a busy city environment.\n\nIt\'s important to note that sometimes there are separate pedestrian signals at intersections, which display "walk" and "don\'t walk" signs to guide pedestrians. If such a pedestrian signal exists in the scene and the walk sign is lit, then pedestrians can safely cross the street while the green traffic light indicates that vehicles should continue moving or turning, if allowed by the traffic sign directions.\n\nIn summary, the presence of a green traffic light at an intersection in the image indicates that vehicles have the right-of-way and can proceed through the intersection. Pedestrians should exercise caution while crossing the street and be aware of the pedestrian signals at the intersection, if available.</s>'], tensor([], size=(0, 480, 640)), tensor([[255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        ...,
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.],
        [255., 255., 255.,  ..., 255., 255., 255.]]), (768, 1024), ['A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human\'s questions. USER: <image>\nWhat color is the traffic light in the image? ASSISTANT: The traffic light in the image is green.</s>USER: Where is the traffic light located in the scene? ASSISTANT: The traffic light is located at a wide intersection in the crosswalk area, within a street intersection that also features a crosswalk and a street sign.</s>USER: Can you describe the surroundings of the traffic light? ASSISTANT: The surroundings of the traffic light include a city street, a tall house sitting between two roads, and various traffic lights and signs. In the background, there are trees that add a touch of greenery to the urban environment.</s>USER: What does the presence of the green traffic light indicate for pedestrians and drivers? ASSISTANT: The green traffic light indicates that it is currently safe for vehicles to pass through the intersection as they have the right-of-way. Pedestrians should be cautious while crossing the street during this time as the green light permits vehicles to move forward. However, this doesn\'t necessarily mean that pedestrians aren\'t allowed to cross, but they should be more vigilant and watch for incoming vehicles, especially in a busy city environment.\n\nIt\'s important to note that sometimes there are separate pedestrian signals at intersections, which display "walk" and "don\'t walk" signs to guide pedestrians. If such a pedestrian signal exists in the scene and the walk sign is lit, then pedestrians can safely cross the street while the green traffic light indicates that vehicles should continue moving or turning, if allowed by the traffic sign directions.\n\nIn summary, the presence of a green traffic light at an intersection in the image indicates that vehicles have the right-of-way and can proceed through the intersection. Pedestrians should exercise caution while crossing the street and be aware of the pedestrian signals at the intersection, if available.</s>'], ['A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human\'s questions. USER: <image>\nWhat color is the traffic light in the image? ASSISTANT: The traffic light in the image is green.</s>USER: Where is the traffic light located in the scene? ASSISTANT: The traffic light is located at a wide intersection in the crosswalk area, within a street intersection that also features a crosswalk and a street sign.</s>USER: Can you describe the surroundings of the traffic light? ASSISTANT: The surroundings of the traffic light include a city street, a tall house sitting between two roads, and various traffic lights and signs. In the background, there are trees that add a touch of greenery to the urban environment.</s>USER: What does the presence of the green traffic light indicate for pedestrians and drivers? ASSISTANT: The green traffic light indicates that it is currently safe for vehicles to pass through the intersection as they have the right-of-way. Pedestrians should be cautious while crossing the street during this time as the green light permits vehicles to move forward. However, this doesn\'t necessarily mean that pedestrians aren\'t allowed to cross, but they should be more vigilant and watch for incoming vehicles, especially in a busy city environment.\n\nIt\'s important to note that sometimes there are separate pedestrian signals at intersections, which display "walk" and "don\'t walk" signs to guide pedestrians. If such a pedestrian signal exists in the scene and the walk sign is lit, then pedestrians can safely cross the street while the green traffic light indicates that vehicles should continue moving or turning, if allowed by the traffic sign directions.\n\nIn summary, the presence of a green traffic light at an intersection in the image indicates that vehicles have the right-of-way and can proceed through the intersection. Pedestrians should exercise caution while crossing the street and be aware of the pedestrian signals at the intersection, if available.</s>'], [{}], False, 'vqa'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/coco/train2017/000000077001.jpg', tensor([[[-1.6384, -1.6384, -1.6213,  ...,  1.9407,  1.9578,  1.9578],
         [-1.6727, -1.6727, -1.6727,  ...,  1.9407,  1.9578,  1.9578],
         [-1.7240, -1.7240, -1.7240,  ...,  1.9578,  1.9578,  1.9578],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.4930, -1.4580, -1.4230,  ...,  2.0959,  2.1134,  2.1134],
         [-1.5105, -1.4930, -1.4755,  ...,  2.0959,  2.1134,  2.1134],
         [-1.5280, -1.5455, -1.5455,  ...,  2.1134,  2.1134,  2.1134],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]],

        [[-1.1770, -1.2641, -1.3687,  ...,  2.4134,  2.4308,  2.4308],
         [-1.3164, -1.3687, -1.4036,  ...,  2.4134,  2.4308,  2.4308],
         [-1.4907, -1.4733, -1.4384,  ...,  2.4134,  2.4134,  2.4134],
         ...,
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000],
         [ 0.0000,  0.0000,  0.0000,  ...,  0.0000,  0.0000,  0.0000]]]), tensor([[[1.6384, 1.6238, 1.6092,  ..., 1.7260, 1.7260, 1.7260],
         [1.6384, 1.6384, 1.6384,  ..., 1.7406, 1.7406, 1.7552],
         [1.6530, 1.6530, 1.6384,  ..., 1.7552, 1.7406, 1.7406],
         ...,
         [0.4559, 0.3683, 0.4413,  ..., 0.8647, 0.8501, 0.8647],
         [0.6019, 0.6603, 0.5727,  ..., 0.9668, 0.9376, 0.8938],
         [0.0763, 0.0909, 0.3683,  ..., 0.8792, 0.8792, 0.9084]],

        [[1.7897, 1.7897, 1.7897,  ..., 1.8798, 1.8798, 1.8798],
         [1.7897, 1.7897, 1.7897,  ..., 1.8798, 1.8798, 1.8948],
         [1.8047, 1.8047, 1.7897,  ..., 1.8798, 1.8798, 1.8798],
         ...,
         [0.5741, 0.5141, 0.5291,  ..., 1.0844, 1.0844, 1.0844],
         [0.7092, 0.7992, 0.7092,  ..., 1.1594, 1.1144, 1.0544],
         [0.1689, 0.1689, 0.5141,  ..., 1.0694, 1.0844, 1.1144]],

        [[1.9468, 1.9468, 1.9468,  ..., 1.9895, 1.9895, 1.9895],
         [1.9468, 1.9468, 1.9468,  ..., 1.9895, 1.9895, 2.0037],
         [1.9610, 1.9610, 1.9468,  ..., 1.9895, 1.9895, 1.9895],
         ...,
         [0.5248, 0.5390, 0.5532,  ..., 1.3211, 1.2927, 1.2927],
         [0.7523, 0.7808, 0.7097,  ..., 1.3496, 1.3069, 1.2500],
         [0.2262, 0.2404, 0.4964,  ..., 1.2927, 1.3069, 1.3354]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the person shown in this image. ASSISTANT: Here's where the person appears in the image.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is tree in this image? Please elaborate your answer and explain why. ASSISTANT: This is the selected tree.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease point out the sea in this picture. ASSISTANT: Take a look at the sea here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease segment the sky in this image. ASSISTANT: Result: sky.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you isolate the boat in this picture? ASSISTANT: The region corresponding to the boat is shown.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         [ True,  True,  True,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True],
         [ True,  True,  True,  ...,  True,  True,  True]],

        [[False, False, False,  ...,  True,  True,  True],
         [False, False, False,  ...,  True,  True,  True],
         [False, False, False,  ...,  True,  True,  True],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[168, 168, 168,  ..., 156, 156, 156],
        [168, 168, 168,  ..., 156, 156, 156],
        [168, 168, 168,  ..., 156, 156, 156],
        ...,
        [154, 154, 154,  ..., 154, 154, 154],
        [154, 154, 154,  ..., 154, 154, 154],
        [154, 154, 154,  ..., 154, 154, 154]]), (683, 1024), ['<image>\nDisplay a segmentation mask for the person shown in this image.', '<image>\nWhat is tree in this image? Please elaborate your answer and explain why.', '<image>\nPlease point out the sea in this picture.', '<image>\nPlease segment the sky in this image.', '<image>\nCould you isolate the boat in this picture?'], ['person', 'tree', 'sea', 'sky', 'boat'], [{'person': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'tree': tensor([[ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        [ True,  True,  True,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'sea': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True],
        [ True,  True,  True,  ...,  True,  True,  True]])}, {'sky': tensor([[False, False, False,  ...,  True,  True,  True],
        [False, False, False,  ...,  True,  True,  True],
        [False, False, False,  ...,  True,  True,  True],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'boat': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg'), ('/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset/ade20k/images/training/ADE_train_00018837.jpg', tensor([[[ 0.6392,  0.6392,  0.6392,  ...,  1.5125,  1.5125,  1.5125],
         [ 0.6392,  0.6392,  0.6392,  ...,  1.5125,  1.5125,  1.5125],
         [ 0.6392,  0.6392,  0.6392,  ...,  1.5125,  1.5125,  1.5125],
         ...,
         [-0.1657, -0.1657, -0.1657,  ...,  0.4337,  0.4337,  0.4337],
         [-0.1657, -0.1657, -0.1657,  ...,  0.4508,  0.4508,  0.4508],
         [-0.1657, -0.1657, -0.1657,  ...,  0.4508,  0.4508,  0.4508]],

        [[ 0.7479,  0.7479,  0.7479,  ...,  1.8158,  1.8158,  1.8158],
         [ 0.7479,  0.7479,  0.7479,  ...,  1.8158,  1.8158,  1.8158],
         [ 0.7479,  0.7479,  0.7479,  ...,  1.8158,  1.8158,  1.8158],
         ...,
         [-0.3025, -0.3025, -0.3025,  ...,  0.3978,  0.3978,  0.3978],
         [-0.3025, -0.3025, -0.3025,  ...,  0.4153,  0.4153,  0.4153],
         [-0.3025, -0.3025, -0.3025,  ...,  0.4153,  0.4153,  0.4153]],

        [[ 1.3502,  1.3502,  1.3502,  ...,  2.3960,  2.3960,  2.3960],
         [ 1.3502,  1.3502,  1.3502,  ...,  2.3960,  2.3960,  2.3960],
         [ 1.3502,  1.3502,  1.3502,  ...,  2.3960,  2.3960,  2.3960],
         ...,
         [-0.0267, -0.0267, -0.0267,  ...,  0.6356,  0.6356,  0.6356],
         [-0.0267, -0.0267, -0.0267,  ...,  0.6531,  0.6531,  0.6531],
         [-0.0267, -0.0267, -0.0267,  ...,  0.6531,  0.6531,  0.6531]]]), tensor([[[ 0.5581,  0.5873,  0.6311,  ...,  1.3172,  1.3026,  1.3026],
         [ 0.5727,  0.6019,  0.6165,  ...,  1.3172,  1.3026,  1.3026],
         [ 0.6165,  0.6311,  0.6311,  ...,  1.3172,  1.3026,  1.3026],
         ...,
         [-0.1134, -0.1134, -0.1134,  ...,  0.2661,  0.2661,  0.2661],
         [-0.1280, -0.1280, -0.1280,  ...,  0.2807,  0.2807,  0.2807],
         [-0.1280, -0.1280, -0.1280,  ...,  0.3683,  0.3683,  0.3829]],

        [[ 0.6341,  0.6642,  0.7092,  ...,  1.5646,  1.5496,  1.5496],
         [ 0.6491,  0.6792,  0.6942,  ...,  1.5646,  1.5496,  1.5496],
         [ 0.6942,  0.7092,  0.7092,  ...,  1.5646,  1.5496,  1.5496],
         ...,
         [-0.2513, -0.2513, -0.2513,  ...,  0.2139,  0.2139,  0.2139],
         [-0.2663, -0.2663, -0.2663,  ...,  0.2289,  0.2289,  0.2289],
         [-0.2663, -0.2663, -0.2663,  ...,  0.3190,  0.3190,  0.3340]],

        [[ 1.0936,  1.1221,  1.1647,  ...,  1.9610,  1.9468,  1.9468],
         [ 1.1078,  1.1363,  1.1505,  ...,  1.9610,  1.9468,  1.9468],
         [ 1.1505,  1.1647,  1.1647,  ...,  1.9610,  1.9468,  1.9468],
         ...,
         [-0.0156, -0.0156, -0.0156,  ...,  0.3968,  0.3968,  0.3968],
         [-0.0298, -0.0298, -0.0298,  ...,  0.4110,  0.4110,  0.4110],
         [-0.0298, -0.0298, -0.0298,  ...,  0.4964,  0.4964,  0.5106]]]), ["A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nPlease highlight the tree in this image. ASSISTANT: Take a look at the tree here.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nWhat is road in this image? Please give some explanation. ASSISTANT: The road portion is displayed below.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the mountain in this picture? ASSISTANT: Sure — this shows the mountain.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nDisplay a segmentation mask for the building shown in this image. ASSISTANT: You can see the building in this frame.</s>", "A chat between a curious human and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the human's questions. USER: <image>\nCould you identify and segment out the sky in this picture? ASSISTANT: Here's where the sky appears in the image.</s>"], tensor([[[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False,  True,  ..., False, False, False],
         [False, False,  True,  ..., False, False, False],
         [False, False,  True,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]],

        [[False, False, False,  ..., False, False, False],
         [False, False, False,  ...,  True,  True,  True],
         [False, False,  True,  ...,  True,  True,  True],
         ...,
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False],
         [False, False, False,  ..., False, False, False]]]), tensor([[255, 255, 255,  ..., 255, 255, 255],
        [255, 255, 255,  ...,   2,   2,   2],
        [255, 255,   2,  ...,   2,   2,   2],
        ...,
        [255, 255,   6,  ...,  11,  11,  11],
        [255, 255,   6,  ...,  11,  11,  11],
        [255, 255,   6,  ...,  11,  11,  11]]), (1024, 1024), ['<image>\nPlease highlight the tree in this image.', '<image>\nWhat is road in this image? Please give some explanation.', '<image>\nCould you identify and segment out the mountain in this picture?', '<image>\nDisplay a segmentation mask for the building shown in this image.', '<image>\nCould you identify and segment out the sky in this picture?'], ['tree', 'road', 'mountain', 'building', 'sky'], [{'tree': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'road': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False,  True,  ..., False, False, False],
        [False, False,  True,  ..., False, False, False],
        [False, False,  True,  ..., False, False, False]])}, {'mountain': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'building': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}, {'sky': tensor([[False, False, False,  ..., False, False, False],
        [False, False, False,  ...,  True,  True,  True],
        [False, False,  True,  ...,  True,  True,  True],
        ...,
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False],
        [False, False, False,  ..., False, False, False]])}], False, 'sem_seg')]
wandb: \ 0.015 MB of 0.015 MB uploaded
wandb: | 0.015 MB of 0.015 MB uploaded
wandb: / 0.015 MB of 0.015 MB uploaded
wandb: - 0.015 MB of 0.015 MB uploaded
wandb: \ 0.015 MB of 0.015 MB uploaded
wandb: | 0.015 MB of 0.015 MB uploaded
wandb: 🚀 View run plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025_bidirbio_2048_maxlen512_epochs25_bsz6_lr0.0003_bidir_bio_feedback_loop_train_prompt_enc_srates_9_7_7_1 at: https://wandb.ai/uiucnlp/plum-training/runs/ltu7jgbz
wandb: ️⚡ View job at https://wandb.ai/uiucnlp/plum-training/jobs/QXJ0aWZhY3RDb2xsZWN0aW9uOjU2ODczNTExNA==/version_details/v68
wandb: Synced 6 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)
wandb: Find logs at: ./wandb/run-20251022_160613-ltu7jgbz/logs
[2025-10-22 16:11:30,291] [INFO] [launch.py:315:sigkill_handler] Killing subprocess 3091899
[2025-10-22 16:11:31,121] [INFO] [launch.py:315:sigkill_handler] Killing subprocess 3091900
[2025-10-22 16:11:31,121] [ERROR] [launch.py:321:sigkill_handler] ['/shared/nas/data/m1/jk100/.conda/envs/llava/bin/python', '-u', 'plum_train_ds.py', '--local_rank=1', '--version=liuhaotian/llava-llama-2-13b-chat-lightning-preview', '--dataset_dir=/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/dataset', '--vision_pretrained=/shared/nas/data/m1/jk100/code/OpenAttrLibrary/LISA/weights/sam_vit_h_4b8939.pth', '--dataset=sem_seg||refer_seg||vqa||reason_seg', '--val_dataset=reason_seg|val', '--sample_rates=9,7,7,1', '--batch_size=6', '--grad_accumulation_steps', '10', '--use_bidir_bio', '--use_feedback_loop', '--ce_loss_weight=1.0', '--dice_loss_weight=8', '--bce_loss_weight=2.0', '--kld_loss_weight=0.1', '--seg_cls_loss_weight=2', '--exp_name=plum-13b_kld_0.1_focal_tversky_8_v1_0shot_w_reasonseg_10222025', '--precision=bf16', '--model_max_length', '512', '--dice_scale_factor', '1000.0', '--epochs', '25', '--train_mask_prompt_encoder', '--focal_tversky_alpha=0.7', '--focal_tversky_beta=0.3', '--bidir_dim_feedforward', '2048', '--auto_resume'] exits with return code = 1
