Rob,

In this case, I'm using UDEV rules to create symlinks at `/dev/tape/...` based on the device wwids/serials, so I don't think it will change assignments.  But I can certainly see why that might be a concern.

SUBSYSTEM=="scsi_generic", ENV{SCSI_TYPE}=="medium changer", ENV{ID_VENDOR}=="SPECTRA", ENV{ID_MODEL}=="PYTHON", ENV{ID_SERIAL_SHORT}=="5000e111ecf150b8", SYMLINK+="tape/changer" SUBSYSTEM=="scsi_tape", KERNEL=="nst[0-9]", ATTRS{wwid}=="naa.5000e111ecf150b5", SYMLINK+="tape/drive_1" SUBSYSTEM=="scsi_generic", KERNEL=="sg*", ATTRS{wwid}=="naa.5000e111ecf150b5", SYMLINK+="tape/drive_1_generic" SUBSYSTEM=="scsi_tape", KERNEL=="nst[0-9]", ATTRS{wwid}=="naa.5000e111ecf150c9", SYMLINK+="tape/drive_2" SUBSYSTEM=="scsi_generic", KERNEL=="sg*",  ATTRS{wwid}=="naa.5000e111ecf150c9", SYMLINK+="tape/drive_2_generic"

[root@granite1 ~]# ls -l /dev/tape/
total 0
drwxr-xr-x 2 root root 260 Oct  2 10:01 by-id
drwxr-xr-x 2 root root 140 Oct  2 10:01 by-path
lrwxrwxrwx 1 root root   7 Oct  2 10:01 changer -> ../sg16
lrwxrwxrwx 1 root root   7 Oct  2 10:01 drive_1 -> ../nst1
lrwxrwxrwx 1 root root   7 Oct  2 10:01 drive_1_generic -> ../sg15
lrwxrwxrwx 1 root root   7 Oct  2 10:01 drive_2 -> ../nst0
lrwxrwxrwx 1 root root   7 Oct  2 10:01 drive_2_generic -> ../sg13
[root@granite1 ~]#

[root@granite1 ~]# ls -l /dev/tape/by-id/
total 0
lrwxrwxrwx 1 root root  9 Oct  2 10:01 scsi-11ECF150B5 -> ../../st1
lrwxrwxrwx 1 root root 10 Oct  2 10:01 scsi-11ECF150B5-nst -> ../../nst1
lrwxrwxrwx 1 root root  9 Oct  2 10:01 scsi-11ECF150C9 -> ../../st0
lrwxrwxrwx 1 root root 10 Oct  2 10:01 scsi-11ECF150C9-nst -> ../../nst0
lrwxrwxrwx 1 root root  9 Oct  2 10:01 scsi-35000e111ecf150b5 -> ../../st1
lrwxrwxrwx 1 root root 10 Oct  2 10:01 scsi-35000e111ecf150b5-nst -> ../../nst1 lrwxrwxrwx 1 root root 10 Oct  2 10:01 scsi-35000e111ecf150b8 -> ../../sg16 lrwxrwxrwx 1 root root 10 Oct  2 10:01 scsi-35000e111ecf150b8-changer -> ../../sg16
lrwxrwxrwx 1 root root  9 Oct  2 10:01 scsi-35000e111ecf150c9 -> ../../st0
lrwxrwxrwx 1 root root 10 Oct  2 10:01 scsi-35000e111ecf150c9-nst -> ../../nst0
lrwxrwxrwx 1 root root 10 Oct  2 10:01 scsi-DE68105842_LL01 -> ../../sg16
[root@granite1 ~]#
[root@granite1 ~]# lsscsi -g
[0:0:67:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sda   /dev/sg0
[0:0:68:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdb   /dev/sg1
[0:0:69:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdd   /dev/sg2
[0:0:70:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdc   /dev/sg3
[0:0:71:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sde   /dev/sg4
[0:0:72:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdf   /dev/sg5
[0:0:73:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdg   /dev/sg6
[0:0:74:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdh   /dev/sg7
[0:0:75:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdi   /dev/sg8
[0:0:76:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdj   /dev/sg9
[0:0:77:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdk   /dev/sg10
[0:0:78:0]   disk    WDC      WUH722012CL5200  AS05  /dev/sdl   /dev/sg11
[0:0:79:0]   enclosu DP       BP_PSV           1.92  -          /dev/sg12
[1:0:0:0]    tape    IBM      ULTRIUM-TDA      SBN0  /dev/st0   /dev/sg13
[1:1:17:0]   enclosu DP       CBL              0.00  -          /dev/sg14
[2:0:0:0]    tape    IBM      ULTRIUM-TDA      SBN0  /dev/st1   /dev/sg15
[2:0:0:1]    mediumx SPECTRA  PYTHON           2.63  /dev/sch0  /dev/sg16
[2:1:17:0]   enclosu DP       CBL              0.00  -          /dev/sg17
[N:0:0:1]    disk    Dell BOSS-N1__1                            -          -
[root@granite1 ~]#



I did work through various btape tests, and they all succeeded without error.  Even in semi-production, these only come up every several weeks, though, so I don't know if that's definitive. Certainly not something I can predict when it will happen.

It's worth noting that, after my initial message, my co-worker found some suspicious looking SAS error messages shortly before the I/O errors in the log.  We don't know yet if that's a cable (already replaced once), SAS card, or drive problem, but we'll continue to pursue it.


Here are the config excerpts requested:

From the director:

Device {
 Name = tapedriveSD1-1
 Device Type = Tape
 Media Type = LTO-10
 Archive Device = /dev/tape/drive_1
 AutomaticMount = yes
 AlwaysOpen = yes
 RemovableMedia = yes
 Random Access = no
 Drive Index = 0

 Maximum Concurrent Jobs = 20

 Spool Directory = /bacula/granite1/spool1
 #Maximum Job Spool Size = 200GB
 Maximum Job Spool Size = 150GB
 Maximum Spool Size = 1900GB

 Autochanger = yes

 Maximum File Size = 32G
 Maximum Block Size = 2M
 Maximum Network Buffer Size = 1000000

 Maximum Changer Wait = 600

}

Device {
 Name = tapedriveSD1-2
 Device Type = Tape
 Media Type = LTO-10
 Archive Device = /dev/tape/drive_2
 AutomaticMount = yes
 AlwaysOpen = yes
 RemovableMedia = yes
 Random Access = no
 Drive Index = 1

 Maximum Concurrent Jobs = 20

 Spool Directory = /bacula/granite1/spool2
 #Maximum Job Spool Size = 200GB
 Maximum Job Spool Size = 150GB
 Maximum Spool Size = 1900GB

 Autochanger = yes

 Maximum File Size = 32G
 Maximum Block Size = 2M
 Maximum Network Buffer Size = 1000000

 Maximum Changer Wait = 600
}

Autochanger {
 Name = SpectraTapeChanger-ctb450
 Device = tapedriveSD1-1
 Device = tapedriveSD1-2
 Changer Device = /dev/tape/changer
 Changer Command = "/usr/lib/bacula/mtx-changer %c %o %S %a %d"

}


And from the SD's config:

Storage {
  Name = granite1-sd
  SDPort = 9103
  WorkingDirectory = "/var/lib/bacula"
  #WorkingDirectory = "/opt/bacula/working"
  #Pid Directory = "/run/bacula"
  Pid Directory = "/var/run"
  Maximum Concurrent Jobs = 75

  TLS Enable = yes
  TLS Require = yes
  TLS CA Certificate File = "/etc/bacula/crt/ca.crt"
  TLS Certificate = "/etc/bacula/crt/granite1.crt"
  TLS Key = "/etc/bacula/crt/granite1.key"

  Heartbeat Interval = 5 minutes
}

Device {
  Name = tapedriveSD1-1
  Device Type = Tape
  Media Type = LTO-10
  Archive Device = /dev/tape/drive_1
  AutomaticMount = yes
  AlwaysOpen = yes
  RemovableMedia = yes
  Random Access = no
  Drive Index = 0

  Maximum Concurrent Jobs = 20

  Spool Directory = /bacula/granite1/spool1
  #Maximum Job Spool Size = 200GB
  Maximum Job Spool Size = 150GB
  Maximum Spool Size = 1900GB

  Autochanger = yes

  Maximum File Size = 32G
  Maximum Block Size = 2M
  Maximum Network Buffer Size = 1000000

  Maximum Changer Wait = 600

}

Device {
  Name = tapedriveSD1-2
  Device Type = Tape
  Media Type = LTO-10
  Archive Device = /dev/tape/drive_2
  AutomaticMount = yes
  AlwaysOpen = yes
  RemovableMedia = yes
  Random Access = no
  Drive Index = 1

  Maximum Concurrent Jobs = 20

  Spool Directory = /bacula/granite1/spool2
  #Maximum Job Spool Size = 200GB
  Maximum Job Spool Size = 150GB
  Maximum Spool Size = 1900GB

  Autochanger = yes

  Maximum File Size = 32G
  Maximum Block Size = 2M
  Maximum Network Buffer Size = 1000000

  Maximum Changer Wait = 600
}

Autochanger {
  Name = SpectraTapeChanger-ctb450
  Device = tapedriveSD1-1
  Device = tapedriveSD1-2
  Changer Device = /dev/tape/changer
  Changer Command = "/usr/lib/bacula/mtx-changer %c %o %S %a %d"

}


On 10/5/26 16:36, Rob Gerber wrote:
Lloyd,

Here is my Level 1 equivalent to Bill's (likely more helpful) email.

I am curious to see the output of the following commands:
ls -lah /dev/tape/by-id/
ls -lah /dev/tape/drive_1
lsscsi

Overall, I suspect "tapedriveSD1-1" (/dev/tape/drive_1) may have had its device assignment change. This is possible, at least. I assume you are familiar with the fact that a hard drive in /dev/sda could change after a reboot. /dev/sda isn't a guaranteed assignment. Your device /dev/tape/drive_1 might be a custom assignment, but I would certainly check it first. Best practice is to give bacula a device path which will not change.

Generally, it is wise to provide a non-rewinding device, to avoid unnecessary rewinds for cases when the tape is already in the drive and ready to write. Excess rewinds needlessly delay backups and accelerate drive mechanism wear.
st = Scsi Tape
nst = Non-rewinding Scsi Tape
If you access a tape drive using an st device (example: /dev/st0) instead of the preferred nst (example: /dev/nst0), you will include an implied rewind at certain times when the drive is accessed. Bacula is fully capable of issuing rewind commands, so we shouldn't do that work for it.


For your tape drive(s) and your autochanger, please share your storage {} and/or autochanger {} definitions from bacula-dir.conf and the relevant device {} and autochanger {} definitions from bacula-sd.conf. Depending on how you set up your system and where you got bacula, the locations and names of these files could vary.

What results do you get when you run the btape tests? *Please note that these tests are DESTRUCTIVE!* Please only run them on tapes which do not contain production data.
https://www.bacula.org/15.0.x-manuals/en/problems/Testing_Your_Tape_Drive_Wit.html#SECTION003210000000000000000

Here is an example of my LTO-8 configuration in bacula 13.x. I only have 1 drive, in a Flexstor II 24 tape changer from Qualstar (manufactured by BDE).
bacula-sd.conf
Device {
  Name = "drive0"
  MediaType = "LTO-8"
  ArchiveDevice = "/dev/tape/by-id/scsi-1234567890-nst"
  RemovableMedia = yes
  RandomAccess = no
  AutomaticMount = yes
  AlwaysOpen = no
  Autochanger = yes
  DriveIndex = 0
  SpoolDirectory = /mnt/spool
  MaximumSpoolSize = 75G
  MaximumFileSize = 75G
  MaximumConcurrentJobs = 1
}
Autochanger {
  Name = "Q24"
  Device = "drive0"
  ChangerDevice = "/dev/tape/by-id/scsi-1BDT_FlexStor_II_0987654321_LL0-changer"
  ChangerCommand = "/opt/bacula/scripts/mtx-changer %c %o %S %a %d"
}

bacula-dir.conf:
Storage {
   Name = "Q24"
   SdPort = 9103
   Address = "my-bacula-host"
   Password = "MyPasswordWhichIHaveRedacted"
   Device = "drive0"
   MediaType = "LTO-8"
   Autochanger = "Q24"
   MaximumConcurrentJobs = 1
}

Regards,
Robert Gerber
402-237-8692 <tel:(402)%20237-8692>
[email protected]


On Mon, Oct 5, 2026 at 4:53 PM Lloyd Brown via Bacula-users <[email protected]> wrote:

    Hey, all.

    I've run Bacula off and on for several years, including v5 and v9,
    both
    using disk for the backup media.  I'm setting up my first instance
    using
    tape, and Bacula 15.  I keep intermittently running into I/O
    errors on
    the tape, which seems to cause Bacula to declare the tape as "Full",
    even if it's only used 10s of MB of a 30TB tape.  Are there any
    general
    recommendations for setup that I might have missed?  Some kind of
    retry
    mechanism?  Some kind of tuning?  I've already replaced SAS
    cables, and
    it doesn't seem to make any real difference.

    In this case, I'm running one RHEL 9.6 host, attached via SAS to 2
    LTO-10 drives, in a Spectra Stack tape library.  Using the stock
    Linux
    tape drivers, not the IBM ITDT specialty driver.  The other tenant in
    the library (Versity) has not had issues even remotely resembling
    these.

    I'm attaching an excerpt from the bacula log when the I/O errors
    occurred.  I'm still trying to find system logs from around the same
    time.  Any recommendations or pointers welcome.

    Thanks,

    Lloyd


-- Lloyd Brown
    HPC Systems Administrator
    Office of Research Computing
    Brigham Young University
    http://rc.byu.edu
    _______________________________________________
    Bacula-users mailing list
    [email protected]
    https://lists.sourceforge.net/lists/listinfo/bacula-users

--
Lloyd Brown
HPC Systems Administrator
Office of Research Computing
Brigham Young University
http://rc.byu.edu
_______________________________________________
Bacula-users mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/bacula-users

Reply via email to