Important fixes (except the new store):
- Fixed a possible use-after-free in the OSD during error handling of initial
commit/rollback of objects in EC pools.
- Fixed a possible free of an invalid pointer in the OSD during read errors
from snapshot/clone chains in EC pools.
- Fixed possibly incorrect handling of commit/rollback operations in EC pools
during pool PG count changes.
- Fixed the inverted fsync enable parameter in the ublk driver (fsync was not
enabled on pools without immediate_commit).
- Added the raw-ls command for debugging purposes to find object versions in
the cluster using listing operations.
New store fixes:
- Improved startup speed by using LSN-based sorting only for objects with a
large number of intermediate versions.
- Added skip_double_claim option as a temporary workaround to fix the rare OSD
startup error with the "double claimed block" message, observed by several
users. This option does not affect data integrity.
- Fixed incorrect rechecking of small writes during startup, which in theory
could lead to duplicate small write object entries on the OSD.
- Fixed fsync operation for disks with a writeback cache (without capacitors):
- Fixed incorrect semantics of consecutive fsyncs (next fsync was not blocked
by the previous one).
- Added fsync when copying small writes from the buffer to the data device
(somehow forgotten during initial development).
- Added fsync after the initial garbage collection during OSD startup.
- Fixed incorrect cast of LSN from uint64 to uint32, breaking fsync when
reaching LSN 2^32.
- Added missing verification of the metadata header checksum during startup.
- Fixed incorrect updating of object checksums in perfect_csum_update=true mode.
- Fixed a possible OSD crash with "assertion failed" when processing a malformed
EC STABILIZE operation.
- Fixed the accounting of active compactor coroutines.
- Removed broken and untested new->old store conversion support.
Minor issues fixed:
- Incorrect accounting of OSD local operation statistics in replicated pools.
- Missing non-zero exitcodes on vitastor-disk resize command errors.
- Missing reset of the list of inconsistent objects during PG restarts.
- Theoretically possible hangs of various OSD operations when working with
completely corrupted objects (without a single available copy), and possibly
in some other very rare situations.
- Incorrect fsyncs when deleting objects from pools without immediate_commit
(on disks with a writeback cache), which previously could leave garbage when
deleting misplaced objects.
- Possible crash/memory corruption of the NFS server during a targeted attack
on NFS-RDMA.
- Possibly incorrect handling of ENOSPC/EIO write errors in replicated pools,
leading to inability to retry the write later.
- Possible crash instead of a clean error exit when starting an OSD with the
old storage engine on a disk with corrupted journal data.
- Shallow copying of PG configuration in the monitor, however, not related to
actual bugs.
- Incorrect checking of allocated blocks in the QEMU driver in an unused code
branch (without the BDRV_WANT_ZERO flag).
- Possible memory leak on read errors of corrupted objects.
- Possible incorrect PG states when corrupted objects are detected.
- Possible failure to mark all "bad" copies of an object during scrubs without
checksums and with a large number of replicas (> 4).
- Incorrect checksum calculation in the old storage engine when
bitmap_granularity < 4096 (a practically unused configuration).
- Theoretically possible OSD crash in rare cases during a scrub and simultaneous
object recovery.
- Theoretically possible OSD crash when handling PING operation errors.
- Slightly suboptimal logic for reusing the RDMA send buffer.
- Possible memory leak when canceling an already running scrub via no_scrub.
- Possible memory corruption when a client (e.g., QEMU code) passes invalid
buffers and the writeback cache is enabled.
- Potentially incorrect search for corrupted parts of EC objects (inability
to find a "good" combination) during a scrub with checksums disabled.
- Possible additional memory usage on the OSD side when handling failed reads
from snapshots (not a leak however - the memory was freed upon client
disconnection).
- Potential sudden write slowdown at certain pg epoch values due to incorrect
epoch update logic in etcd.
Vitastor
The Idea
Make Clustered Block Storage Fast Again.
Vitastor is a distributed block, file and object SDS, direct replacement of Ceph RBD, CephFS and RGW, and also internal SDS's of public clouds. However, in contrast to them, Vitastor is fast and simple at the same time. The only thing is it's slightly young :-).
Vitastor is architecturally similar to Ceph which means strong consistency, primary-replication, symmetric clustering and automatic data distribution over any number of drives of any size with configurable redundancy (replication or erasure codes/XOR).
Vitastor targets primarily SSD and SSD+HDD clusters with at least 10 Gbit/s network, supports TCP and RDMA and may achieve 4 KB read and write latency as low as ~0.1 ms with proper hardware which is ~10 times faster than other popular SDS's like Ceph or internal systems of public clouds.
Vitastor supports QEMU, UBLK, NBD, NFS protocols, OpenStack, OpenNebula, Proxmox, Kubernetes drivers. More drivers may be created easily.
Read more details in the documentation. You can start from here: Quick Start.
Talks and presentations
- KuberConf'2025: video
- Highload'2025: video, youtube, presentation (in Russian, in English)
- Highload'2022: presentation (in Russian), video
- DevOpsConf'2021: presentation (in Russian, in English), video
Documentation
- Introduction
- Installation
- Configuration
- Usage
- vitastor-cli (command-line interface)
- vitastor-disk (disk management tool)
- fio for benchmarks
- UBLK for kernel mounts
- NBD - old interface for kernel mounts
- QEMU, qemu-img and VDUSE
- NFS clustered file system and pseudo-FS proxy
- Administration
- Performance
Author and License
Copyright (c) Vitaliy Filippov (vitalif [at] yourcmc.ru), 2019+
Join Vitastor Telegram Chat: https://t.me/vitastor
All server-side code (OSD, Monitor and so on) is licensed under the terms of Vitastor Network Public License 1.1 (VNPL 1.1), a copyleft license based on GNU GPLv3.0 with the additional "Network Interaction" clause which requires opensourcing all programs directly or indirectly interacting with Vitastor through a computer network and expressly designed to be used in conjunction with it ("Proxy Programs"). Proxy Programs may be made public not only under the terms of the same license, but also under the terms of any GPL-Compatible Free Software License, as listed by the Free Software Foundation. This is a stricter copyleft license than the Affero GPL.
Please note that VNPL doesn't require you to open the code of proprietary software running inside a VM if it's not specially designed to be used with Vitastor.
Basically, you can't use the software in a proprietary environment to provide its functionality to users without opensourcing all intermediary components standing between the user and Vitastor or purchasing a commercial license from the author 😀.
Client libraries (cluster_client and so on) are dual-licensed under the same VNPL 1.1 and also GNU GPL 2.0 or later to allow for compatibility with GPLed software like QEMU and fio.
You can find the full text of VNPL-1.1 in the file VNPL-1.1.txt. GPL 2.0 is also included in this repository as GPL-2.0.txt.