Do all NBD configuration in the child process, after the last fork. Why? It's needed because there is a race condition in the Linux kernel nbd driver in nbd_add_socket() - it saves `current` task pointer as `nbd->task_setup` and then rechecks if the new `current` is the same. Problem is that if that process is already dead, `current` may be freed and then replaced by another process with the same pointer value. So the check passes and NBD allows a different process to set up a device which is already set up. Proper fix would have to be done in the kernel code, but the workaround is obviously to perform NBD setup from the process which will then actually call NBD_DO_IT. That process stays alive during the whole time of NBD device execution and the (nbd->task_setup != current) check always works correctly, and we don't accidentally break previous NBD devices while setting up a new device. Forking to check every device is of course rather slow, so we also do an additional check by calling list_mapped() before searching for a free NBD device.
Vitastor
The Idea
Make Clustered Block Storage Fast Again.
Vitastor is a distributed block and file SDS, direct replacement of Ceph RBD and CephFS, and also internal SDS's of public clouds. However, in contrast to them, Vitastor is fast and simple at the same time. The only thing is it's slightly young :-).
Vitastor is architecturally similar to Ceph which means strong consistency, primary-replication, symmetric clustering and automatic data distribution over any number of drives of any size with configurable redundancy (replication or erasure codes/XOR).
Vitastor targets primarily SSD and SSD+HDD clusters with at least 10 Gbit/s network, supports TCP and RDMA and may achieve 4 KB read and write latency as low as ~0.1 ms with proper hardware which is ~10 times faster than other popular SDS's like Ceph or internal systems of public clouds.
Vitastor supports QEMU, NBD, NFS protocols, OpenStack, OpenNebula, Proxmox, Kubernetes drivers. More drivers may be created easily.
Read more details in the documentation. You can start from here: Quick Start.
Talks and presentations
- DevOpsConf'2021: presentation (in Russian, in English), video
- Highload'2022: presentation (in Russian), video
Documentation
- Introduction
- Installation
- Configuration
- Usage
- vitastor-cli (command-line interface)
- vitastor-disk (disk management tool)
- fio for benchmarks
- NBD for kernel mounts
- QEMU and qemu-img
- NFS clustered file system and pseudo-FS proxy
- Administration
- Performance
Author and License
Copyright (c) Vitaliy Filippov (vitalif [at] yourcmc.ru), 2019+
Join Vitastor Telegram Chat: https://t.me/vitastor
All server-side code (OSD, Monitor and so on) is licensed under the terms of Vitastor Network Public License 1.1 (VNPL 1.1), a copyleft license based on GNU GPLv3.0 with the additional "Network Interaction" clause which requires opensourcing all programs directly or indirectly interacting with Vitastor through a computer network and expressly designed to be used in conjunction with it ("Proxy Programs"). Proxy Programs may be made public not only under the terms of the same license, but also under the terms of any GPL-Compatible Free Software License, as listed by the Free Software Foundation. This is a stricter copyleft license than the Affero GPL.
Please note that VNPL doesn't require you to open the code of proprietary software running inside a VM if it's not specially designed to be used with Vitastor.
Basically, you can't use the software in a proprietary environment to provide its functionality to users without opensourcing all intermediary components standing between the user and Vitastor or purchasing a commercial license from the author 😀.
Client libraries (cluster_client and so on) are dual-licensed under the same VNPL 1.1 and also GNU GPL 2.0 or later to allow for compatibility with GPLed software like QEMU and fio.
You can find the full text of VNPL-1.1 in the file VNPL-1.1.txt. GPL 2.0 is also included in this repository as GPL-2.0.txt.