Laparoscopic surgery requires rigorous training, but recent advances in machine learning may enable automated video-based assessment of surgical skills. However, progress is currently limited due to the lack of large-scale annotated datasets. To overcome this challenge, we introduce the Laparoscopic Skills Analysis and Assessment (LASANA) dataset, which consists of 1,270 stereo video recordings of four basic laparoscopic training tasks. Each recording contains a structured skill rating derived from three independent raters, along with binary labels that identify task-specific errors. The majority of the recordings were captured during a laparoscopic training course, reflecting natural skill variation among participants. To enable robust benchmarking of current and future video-based skill assessment and error recognition methods, Isabel Funke, Sebastian Bodenstedt, Felix von Bechtolsheim, Florian Oehme, Michael Maluschke, Stefanie Herrlich, Jürgen Weitz, Marius Distler, Sören Torge Mees, Stefanie Led by provides predefined data splits for each task and presents baseline model results for comparative analysis. Speidel.
This new resource addresses a critical limitation in the field, the lack of large annotated datasets, which has hindered the advancement of deep learning models for surgical skill assessment.
LASANA consists of 1,270 stereo video recordings capturing four basic laparoscopic training tasks performed by a diverse group of participants. Each video is carefully annotated with a structured skill rating obtained by consensus of three independent expert raters, along with binary labels that identify errors in specific tasks.
The creation of the dataset reflects a commitment to realistic training scenarios, with the majority of recordings coming from actual laparoscopic training courses. This ensures that the data includes the natural variation in skill levels exhibited by trainees and provides a more solid foundation for model development.
To facilitate standardized benchmarks, researchers provided predefined data splits by task, allowing for a fair comparison of existing and new approaches to video-based skill assessment and error recognition. A baseline model has also been implemented and its results published, serving as an important reference point for future research efforts.
LASANA’s comprehensive annotation extends beyond simple skill scores to include detailed error identification and structured skill assessment. These detailed labels allow for a more nuanced analysis of surgical techniques and facilitate the development of models that can provide targeted feedback to trainees.
The size of this dataset is significantly larger than previously available resources, such as the JIGSAWS dataset with 103 videos, ROSMA with 206 videos, and AIxSuture with 314 videos, and is expected to enable new capabilities for automatic evaluation. The average video duration per task ranges from 2 minutes 32 seconds to 4 minutes 30 seconds, providing sufficient data for robust model training.
This study has the potential to transform laparoscopic surgical training by providing an objective, consistent, and cost-effective evaluation method. By automating skill assessment, LASANA paves the way for personalized training programs, improved feedback mechanisms, and ultimately improved surgical performance and patient outcomes. The study collected 1270 recordings performed by 70 participants on four basic laparoscopic training tasks: peg transfer, circle cut, balloon resection, and suture and knot tying.
Recordings were acquired using a Karl Storz TIPCAM 1 S 3D LAP 30° endoscope, providing a synchronized stereo view of the surgical scene within the Laparo Aspire training box. Each participant’s performance was recorded using a Karl Storz laparoscopic instrument, and the left camera video stream was displayed as visual feedback during task completion.
Skills assessment employed a structured assessment inspired by the Global Operative Assessment of Laparoscope Skills (GOALS) tool and was aggregated from three independent raters. Inspired by GOALS, this rating system evaluates performance on five dimensions, assigns scores on a five-point Likert scale, and ultimately yields a total score representing overall skill.
To ensure reliability of the data, the research team used Lin’s concordance correlation coefficient ρc to quantify inter-rater agreement, achieving values above 0.65 for all tasks except circle cutting, resulting in a ρc of 0.49. In addition, the dataset contains binary labels that indicate the presence or absence of task-specific errors, such as falling objects or punctured balloons, enabling the development of complementary error recognition algorithms. Each recording incorporates a structured skill rating derived from the consensus of three independent raters, along with binary labels that identify the presence or absence of task-specific errors.
This dataset primarily contains records from laparoscopic training courses and accurately reflects the varying skill levels of participating trainees. Predefined data partitions are provided to enable benchmarking of existing and novel video-based skill assessment approaches and error recognition algorithms for each task.
A deep learning model was implemented to establish baseline results and serve as a comparison reference point for future studies. The four laparoscopic training tasks included in the dataset are object manipulation, cutting, and suturing exercises commonly used in surgical curricula. The average video duration of the recordings was approximately 2 minutes 32 seconds for the first task, 3 minutes 32 seconds for the second task, 3 minutes 55 seconds for the third task, and 4 minutes 30 seconds for the last task.
The dataset includes annotations detailing experience levels, skill ratings, surgical movements, and task-specific errors, providing a comprehensive resource for analysis. The 70 participants who contributed to the LASANA dataset generated a total of 1,270 videos, of which approximately 314–329 videos were dedicated to each of the four tasks.
Three independent raters provide skill ratings for each video, ensuring a robust and reliable assessment of surgical performance. The availability of both structured skill assessments and binary error labels facilitates the development of models that can both estimate skill level and detect errors. It consists of 1,270 stereo video recordings documenting four basic laparoscopic training tasks, a significant increase in scale compared to existing datasets such as JIGSAWS, ROSMA, and AIxSuture.
Each video is accompanied by a structured skill assessment derived from multiple independent raters and binary labels that identify task-specific errors, providing detailed information for automated analysis. This dataset will facilitate the development and benchmarking of video-based systems designed to automatically assess surgical skills and recognize errors.
The recordings are primarily from laparoscopic training courses and capture the natural changes in performance level that can be expected during skill acquisition. Predefined data splits are included to standardize the evaluation procedure and enable meaningful comparisons between different approaches, with baseline results from deep learning models provided for reference.
Recognizing the current limitations in the data available to train robust assessment models, LASANA provides a valuable resource to the surgical training community. The authors anticipate that this dataset could accelerate advances in automated skill assessment and lead to more objective and efficient training programs. Future work may focus on expanding the dataset to include more complex surgical procedures and diverse participant populations, further increasing the versatility and applicability of automated assessment tools.
