On Thu, Jul 23, 2026 at 2:19 AM Miklos Szeredi <[email protected]> wrote:
>
> On Thu, 23 Jul 2026 at 11:01, Andy Lutomirski <[email protected]> wrote:
>
> > - O_XYZ flags to open() are all kinds of awful, for historical reasons that 
> > are not fundamental to the concept.
>
> I'm fine with limiting this to openat2().
>
> > - symlinks. If we have a new API where opening /proc/something/magic/blah 
> > can access something that ought to be inaccessible when accessed 
> > intentionally and with privilege, a symlink pointing at /proc/something/… 
> > can cause the API to be used inadvertently.  Admittedly we have this 
> > problem with basically all symlinks, so this isn’t exactly unique.
> >
> > - our nasty fs permission model. We gave a sort of gnarly mix of a bit of 
> > fd-based permission and mostly mode/ACL-based permissions for path 
> > traversal and opening, and mapping this nicely only new APIs (as opposed to 
> > actual files and directories) can have unpleasant results.
>
> Pseudo filesystems don't usually do mode/acl based permission checks.
>  I don't see any issue with that.
>
> > Of course, I’m busy arguing (slowly and without a concrete proposal) that 
> > we should have proper capability-like fds, and maybe that’s kind of an 
> > answer to this:
> >
> > What if we had an API to get an fd to the “control filesystem” for a 
> > superblock, like your O_ALT but as a real syscall or maybe only accessible 
> > via one of the newer and less janky open variants?  And what if the 
> > resulting fd and the filesystem tree it represented had a few properties 
> > that made it very different from normal directory fds:
> >
> > - You cannot mount anything on it or its subdirectories, nor can you 
> > open_tree or otherwise mount it anywhere.  But you can open_tree the very 
> > specific things in it that point outside of the special API (e.g. the 
> > overlayfs layers).
>
> This patch provides exeactly those properties.
>
> > - Privilege is fully captured by the original call that gets you the fd.  
> > current->cred is not checked when *using* it except to the extent that you 
> > might need privileges over your own namespaces to do operations that might 
> > affect them.
>
> This is trivial to add to metafs: just store a ref to current->cred
> besides the path.

We have f_cred that might be good enough, although I'm not sure it
propagates through path lookup.

>
> Not sure how this security model would work, though.

If I have CAP_SYS_ADMIN and I open the metafs and then drop privilege,
then I would still be able to do all the metafs operations as if I had
CAP_SYS_ADMIN so long as I use the fd I opened to do those operations.

>
> >
> > - You can’t fchdir or (hypothetically) fchroot into it.  (Not sure how 
> > important this is.)
>
> Not sure if there are internal mounts that have directories?  If not,
> we can just add a check to fchdir/chroot against mnt_ns being NULL.
>
> > - Maybe you can’t follow /proc/pid/fd/N links into it either?  Or maybe 
> > that would break CRIU too badly.
> >
> > The basic idea here is to try to treat it like an API that happens to use 
> > the open machinery but not as part of the filesystem hierarchy.
>
> I'm not opposed to limiting this in various ways, though I don't
> really see the advantage.

I'll sleep on it.  Maybe it's all fine.

Reply via email to