Virtual intraoperative CT (viCT): sequential anatomic updates for modeling tissue resection throughout endoscopic sinus surgery

3D reconstruction

Metrically scaled intraoperative geometry was reconstructed from monocular endoscopic video acquired with a 4-mm rigid 30° endoscope with a 1080p video tower (Karl Storz, Tuttlingen, Germany), and exported for processing via a flash drive. Video was recorded at 1080p resolution and 30 frames/s. At each surgical interval, a 3-5 second video sweep was acquired to view the resection cavity from multiple overlapping positions and orientations. After removing frames containing blur, instrument occlusion, irrigation, or severe specular reflection, the remaining frames were processed using a depth-supervised Neural Radiance Field (NeRF) framework (Fig. 1), producing a spatially consistent 3D representation for CT-grid voxelization. [4]

Camera intrinsics \(\textbf\) and poses \(\textbf_i=[\textbf_i|\textbf_i]\in SE(3)\) were estimated with COLMAP, with projection

$$\begin \textbf_i \sim \textbf(\textbf_i\textbf+\textbf_i). \end$$

(1)

The NeRF represents radiance and density as

$$\begin F_:(\textbf,\textbf)\rightarrow (\sigma ,\textbf), \end$$

(2)

where \(\textbf\in \mathbb ^3\) and \(\textbf\) is viewing direction. Following EndoPerfect, spatial positions are encoded using multi-resolution hash encoding \(\phi (\textbf)\) and view dependence using spherical harmonics

$$\begin \gamma (\textbf)= \big [ Y_0^0(\textbf),Y_1^(\textbf),\ldots ,Y_}^}(\textbf) \big ]. \end$$

(3)

The model is trained by minimizing photometric reconstruction error

$$\begin \mathcal _}=\sum _\Vert I_i-\hat_i\Vert _2^2, \end$$

(4)

with \(\hat_i\) obtained through discrete volume rendering. For a ray \(\textbf(t)=\textbf+t\textbf\) sampled at \(\_^\),

$$\begin \hat}(\textbf)= \sum _^ T_k(1e^)\,\textbf_k, \qquad T_k\exp \!\left( -\sum _\sigma _j\delta _j\right) ,\nonumber \\ \end$$

(5)

where \(\sigma _k=\sigma (\textbf(t_k))\), \(\textbf_k=\textbf(\textbf(t_k),\textbf)\), and \(\delta _k=t_-t_k\).

Metric depth is obtained by synthesizing a virtual stereo partner view within the trained NeRF. For each training pose \(\textbf_}\), a novel pose \(\textbf_}\) is optimized to maximize stereo similarity via zero-mean normalized cross-correlation (ZNCC) subject to fixed baseline b and coplanar motion:

$$\begin \textbf_}^ = \arg \max __}\in \mathcal } \textrm\!\left( I_}, I_}(\textbf_})\right) . \end$$

(6)

The rendered pair \((I_},I_})\) is used for stereo depth estimation, and depth maps are fused across views to produce a dense metrically scaled reconstruction for CT-grid integration. [4, 7, 18]

Fig. 1Fig. 1

Flowchart of simulated stereoscopy NeRF-based reconstruction workflow

Registration

To generate viCT volumes, each intraoperative 3D reconstruction was rigidly registered to the corresponding preoperative CT (pCT). Semi-automatic, landmark-based rigid registration was performed in 3D Slicer using anatomically corresponding fiducials visible in both datasets. Landmarks were selected from stable bony structures, placed by researchers, and verified by an otolaryngologist to ensure anatomical consistency.

Between 3–10 landmarks were used per specimen and surgical interval. Common fiducials included the nasofrontal beak, middle turbinate axilla, junction of the maxillary sinus roof with the lamina papyracea in the plane of the posterior maxillary wall, sphenoid sinus landmarks (ostium, rostrum, floor, posterior wall), orbital entry of the anterior ethmoid artery, cribriform lateral lamella/fovea ethmoidalis junction, and basal lamella/skull base attachment.

Virtual intraoperative CT (viCT) updatingFig. 2Fig. 2

Ray-based viCT updating workflow and example of tissue resection modeling

The viCT update is computed directly in the native preoperative CT (pCT) voxel grid using DICOM metadata (origin, spacing, orientation). Let the pCT be \(I_}:\Omega \subset \mathbb ^3\rightarrow \mathbb \) with voxel index \(\textbf=(i_x,i_y,i_z)\). The corresponding physical coordinate \(\textbf\in \mathbb ^3\) is

$$\begin \textbf(\textbf) = \textbf_} + \textbf_} \begin s_x i_x\\ s_y i_y\\ s_z i_z \end, \end$$

(7)

where \(\textbf_}\) is the DICOM origin, \(\textbf_}\in SO(3)\) the direction cosine matrix, and \((s_x,s_y,s_z)\) the voxel spacings (mm). All reconstruction-derived geometry is rigidly registered in this coordinate frame.

A binary anatomical mask is generated from the pCT by thresholding

$$\begin M_}(\textbf) = \mathbb \!\left( I_}(\textbf) > \tau \right) , \end$$

(8)

with threshold \(\tau \) (here, \(\tau =-300\) HU). The original HU-valued pCT volume is preserved; only voxel occupancy is modified.

Intraoperative reconstructions are exported as an STL mesh with an associated fiducial point (FCSV) defining the reconstructed camera origin \(\textbf_c\). Because STL meshes may be hollow, a volumetric occupancy mask is generated by ray-based voxelization. Let \(\_m\}_^\) denote mesh vertices. For each vertex,

$$\begin \textbf_m=\frac_m-\textbf_c}_m-\textbf_c\Vert }, \qquad \ell _m=\Vert \textbf_m-\textbf_c\Vert . \end$$

(9)

Rays are parameterized as

$$\begin \textbf_m(t)=\textbf_c+t\textbf_m, \qquad t\in [0,\ell _m]. \end$$

(10)

Sampled points along each ray are mapped to voxel indices using the inverse of Eq. (7), forming a sparse occupancy scaffold. Binary dilation, morphological closing, and hole filling produce a watertight volumetric mask \(M_}:\Omega \rightarrow \\) expressed in the pCT grid.

Voxels in the pCT mask that overlap the reconstructed free-space mask, and the retained anatomical occupancy used to generate viCT are defined as:

$$\begin M_} = M_} \wedge M_}, \qquad M_} = M_} \wedge \lnot M_}. \end$$

(11)

and the final viCT image is obtained from the original HU-valued pCT volume as

$$\begin I_}(\textbf) = I_}(\textbf), & M_}(\textbf)=1,\\ I_}, & \text , \end\right. } \end$$

(12)

where \(I_}\) is an air-equivalent HU value (e.g., \(-1000\) HU). Unchanged anatomy retains original HU intensities and DICOM metadata, while resected tissue is removed via intensity reassignment, yielding a CT-format volume in the native pCT coordinate system.

For quantitative validation against interval CT, binary masks derived from \(I_}\) and ground-truth CT are compared within a reconstruction-defined ROI. Surface distances computed in voxel units are converted to millimeters using mean voxel spacing \(\bar=\tfrac(s_x+s_y+s_z)\):

$$\begin d_} = d_}\bar. \end$$

(13)

This process can be repeated as additional reconstructions are generated, enabling sequential HU-preserving CT updates throughout surgery (Fig. 2).

Anatomy-specific analysis

To evaluate whether global overlap metrics obscured localized reconstruction failure, individual STL reconstructions were grouped by the anatomical region represented (anterior ethmoid, maxillary, middle turbinate, nasal floor, posterior ethmoid, or sphenoid). Because individual reconstructions were designed to represent local portions of the surgical field rather than complete surrounding anatomy, evaluation was restricted to voxels within a reconstruction-derived ROI. Predicted surgical change via the viCT was defined relative to the preoperative CT and compared with change observed between preoperative and interval ground-truth CT, thereby excluding unchanged anatomy from the primary comparison. Sensitivity, positive predictive value (PPV), DSC, Jaccard index, ASSD, RMSD, and HD95 were calculated for each reconstruction and summarized by anatomical region. 104 3D reconstructions were labeled with anatomical-region and were included in the region-specific analysis.

Comments (0)

No login
gif